makaron-cli 0.16.0 → 0.16.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/README.md +13 -1
- package/bin/makaron.mjs +47 -3
- package/package.json +1 -1
- package/skills/makaron/SKILL.md +19 -1
package/README.md
CHANGED
|
@@ -107,7 +107,7 @@ npx makaron-cli chat --project auto --image photo.jpg --json -b "make it cinemat
|
|
|
107
107
|
npx makaron-cli chat --project auto --image img1.jpg --image img2.jpg --json -b "combine these"
|
|
108
108
|
```
|
|
109
109
|
|
|
110
|
-
`chat` routes image and video models automatically, but you may select the Agent LLM with `--agent-model`. Accepted values are `auto`, the base model IDs (`gpt-6-luna`, `gpt-6-sol`, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-5.6-luna`, `grok-4.6`, `deepseek-v4-pro`, `deepseek-flash`), and the personal-plan routes (`gpt-6-luna-codex-subscription`, `gpt-6-sol-codex-subscription`, `gpt-5.6-terra-codex-subscription`, `gpt-5.6-sol-codex-subscription`, `gpt-5.6-luna-codex-subscription`, `grok-4.6-grok-subscription`). `auto` resolves to GPT-6 Luna through the Codex subscription for eligible accounts (including admins) and through Azure API otherwise. Base GPT-5.6 and GPT-6 IDs select Azure API and base `grok-4.6` selects OpenRouter API; the suffixed IDs explicitly request the corresponding personal plan where authorized. This flag changes only the reasoning/tool-calling Agent LLM.
|
|
110
|
+
`chat` routes image and video models automatically, but you may select the Agent LLM with `--agent-model`. Accepted values are `auto`, the base model IDs (`gpt-6-luna`, `gpt-6.1-sol`, `gpt-6-sol`, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-5.6-luna`, `grok-4.6`, `deepseek-v4-pro`, `deepseek-flash`), and the personal-plan routes (`gpt-6-luna-codex-subscription`, `gpt-6.1-sol-codex-subscription`, `gpt-6-sol-codex-subscription`, `gpt-5.6-terra-codex-subscription`, `gpt-5.6-sol-codex-subscription`, `gpt-5.6-luna-codex-subscription`, `grok-4.6-grok-subscription`). `auto` resolves to GPT-6 Luna through the Codex subscription for eligible accounts (including admins) and through Azure API otherwise. Base GPT-5.6, GPT-6 and GPT-6.1 IDs select Azure API and base `grok-4.6` selects OpenRouter API; the suffixed IDs explicitly request the corresponding personal plan where authorized. This flag changes only the reasoning/tool-calling Agent LLM.
|
|
111
111
|
|
|
112
112
|
```bash
|
|
113
113
|
# Explicit lower-cost Agent LLM for a controlled comparison
|
|
@@ -417,8 +417,20 @@ npx makaron-cli video create --script "continue the camera move into the next be
|
|
|
417
417
|
|
|
418
418
|
# 4. Check status
|
|
419
419
|
npx makaron-cli video status <taskId>
|
|
420
|
+
|
|
421
|
+
# Rebuild only 2–5 seconds and publish the complete video into the project.
|
|
422
|
+
npx makaron-cli video retake --video input.mp4 --start 2 --end 5 --prompt "Make the umbrella red" --project <id> --model seedance-2.5 --wait --json
|
|
423
|
+
```
|
|
424
|
+
|
|
425
|
+
Video segment editing also works directly from natural-language chat, without a GUI selection:
|
|
426
|
+
|
|
427
|
+
```bash
|
|
428
|
+
npx makaron-cli chat --project <id> "把 @1 的 18–21 秒切成多机位,其余不变"
|
|
429
|
+
npx makaron-cli chat --project <id> "把 @1 的最后三秒改成夜景"
|
|
420
430
|
```
|
|
421
431
|
|
|
432
|
+
Chat resolves the source/range, inspects the footage, expands the edit instruction, and delivers the complete video. It asks only when the source or scope remains ambiguous. Precise cuts, subtitles, dubbing and extension also start from chat. Local editing accepts a 0.1–15s range in a source up to 120s; longer scopes use the appropriate whole-video workflow rather than silently shortening the request.
|
|
433
|
+
|
|
422
434
|
For project/timeline video editing, use:
|
|
423
435
|
|
|
424
436
|
```bash
|
package/bin/makaron.mjs
CHANGED
|
@@ -32,11 +32,13 @@ const CHAT_AGENT_MODELS = [
|
|
|
32
32
|
'auto',
|
|
33
33
|
'gpt-6-luna',
|
|
34
34
|
'gpt-6-sol',
|
|
35
|
+
'gpt-6.1-sol',
|
|
35
36
|
'gpt-5.6-terra',
|
|
36
37
|
'gpt-5.6-sol',
|
|
37
38
|
'gpt-5.6-luna',
|
|
38
39
|
'gpt-6-luna-codex-subscription',
|
|
39
40
|
'gpt-6-sol-codex-subscription',
|
|
41
|
+
'gpt-6.1-sol-codex-subscription',
|
|
40
42
|
'gpt-5.6-terra-codex-subscription',
|
|
41
43
|
'gpt-5.6-sol-codex-subscription',
|
|
42
44
|
'gpt-5.6-luna-codex-subscription',
|
|
@@ -512,9 +514,9 @@ Options:
|
|
|
512
514
|
--audio <file|url> Attach a song, beat, or voice reference. MP3/WAV, repeatable.
|
|
513
515
|
--media-manifest <file|-> Import typed image/video media before this run.
|
|
514
516
|
--skill <id|label|name> Use an installed skill or auto-install a matched marketplace skill.
|
|
515
|
-
--agent-model <id> Agent LLM only: auto, gpt-6-luna, gpt-6-sol, gpt-5.6-terra, gpt-5.6-sol,
|
|
517
|
+
--agent-model <id> Agent LLM only: auto, gpt-6-luna, gpt-6.1-sol, gpt-6-sol, gpt-5.6-terra, gpt-5.6-sol,
|
|
516
518
|
gpt-5.6-luna, grok-4.6, deepseek-v4-pro, deepseek-flash, or a
|
|
517
|
-
gpt-6-*-codex-subscription, gpt-5.6-*-codex-subscription or
|
|
519
|
+
gpt-6.1-sol-codex-subscription, gpt-6-*-codex-subscription, gpt-5.6-*-codex-subscription or
|
|
518
520
|
grok-4.6-grok-subscription personal-plan route.
|
|
519
521
|
--background, -b Submit and print a runId.
|
|
520
522
|
--json Output structured JSON (includes per-run "usage" credits).
|
|
@@ -2079,7 +2081,7 @@ Commands:
|
|
|
2079
2081
|
|
|
2080
2082
|
edit [--image <file>] "prompt" Image edit, text-to-image, or transparent PNG
|
|
2081
2083
|
analyze --video <file|url> Analyze video content
|
|
2082
|
-
video script|create|status
|
|
2084
|
+
video script|create|retake|status H3 Max, Wan, Seedance, Grok, lip-sync; interval retake
|
|
2083
2085
|
music create|status Music generation
|
|
2084
2086
|
|
|
2085
2087
|
admin Admin commands (skills, credits, upload, set-admin)
|
|
@@ -2232,6 +2234,7 @@ Not sure which built-in skill to use? Start with:
|
|
|
2232
2234
|
} else if (topic === 'video') {
|
|
2233
2235
|
if (subtopic === 'script') console.log('Usage: makaron video script --image <file> [--image <file>] [--lang en|zh] "direction"');
|
|
2234
2236
|
else if (subtopic === 'create') printVideoCreateHelp();
|
|
2237
|
+
else if (subtopic === 'retake') console.log('Usage: makaron video retake --video <url|file> --start <seconds> --end <seconds> --prompt "change" [--model seedance-2.5-eco|seedance-2.5|fal-h3-max] [--project <id>] [--wait] [--json] [--request-id <uuid>] [--audio-mode original|generated]');
|
|
2235
2238
|
else if (subtopic === 'status') console.log('Usage: makaron video status <taskId> | --snapshot <snapshotId> [--wait]');
|
|
2236
2239
|
else printVideoHelp();
|
|
2237
2240
|
} else if (topic === 'music') {
|
|
@@ -3085,6 +3088,47 @@ if (!command || command === '--help' || command === '-h' || command === 'help')
|
|
|
3085
3088
|
const text = result?.content?.find(c => c.type === 'text')?.text;
|
|
3086
3089
|
if (text) console.log(text);
|
|
3087
3090
|
|
|
3091
|
+
} else if (sub === 'retake') {
|
|
3092
|
+
const params = { model: 'fal-h3-max' };
|
|
3093
|
+
let wait = false, json = false;
|
|
3094
|
+
for (let i = 2; i < args.length; i++) {
|
|
3095
|
+
const flag = args[i];
|
|
3096
|
+
if (flag === '--wait') wait = true;
|
|
3097
|
+
else if (flag === '--json') json = true;
|
|
3098
|
+
else if (flag === '--video' && args[i + 1]) params.video_url = args[++i];
|
|
3099
|
+
else if (flag === '--start' && args[i + 1]) params.start = Number(args[++i]);
|
|
3100
|
+
else if (flag === '--end' && args[i + 1]) params.end = Number(args[++i]);
|
|
3101
|
+
else if (flag === '--audio-mode' && args[i + 1]) params.audio_mode = args[++i];
|
|
3102
|
+
else if (flag === '--prompt' && args[i + 1]) params.prompt = args[++i];
|
|
3103
|
+
else if ((flag === '--model' || flag === '--video-model') && args[i + 1]) params.model = args[++i];
|
|
3104
|
+
else if (flag === '--project' && args[i + 1]) params.project_id = args[++i];
|
|
3105
|
+
else if (flag === '--request-id' && args[i + 1]) params.request_id = args[++i];
|
|
3106
|
+
else { console.error('Usage: makaron video retake --video <url|file> --start <seconds> --end <seconds> --prompt "change" [--model seedance-2.5-eco|seedance-2.5|fal-h3-max] [--project <id>] [--wait] [--json] [--request-id <uuid>] [--audio-mode original|generated]'); process.exit(1); }
|
|
3107
|
+
}
|
|
3108
|
+
if (!params.video_url || !params.prompt || !Number.isFinite(params.start) || !Number.isFinite(params.end)
|
|
3109
|
+
|| params.start < 0 || params.end - params.start < .1 || params.end - params.start > 15
|
|
3110
|
+
|| (params.audio_mode && !['original','generated'].includes(params.audio_mode))
|
|
3111
|
+
|| !['seedance-2.5-eco', 'seedance-2.5', 'fal-h3-max'].includes(params.model)) {
|
|
3112
|
+
console.error('Retake needs one video, a prompt, a valid 0.1–15 second interval, and a supported model.'); process.exit(1);
|
|
3113
|
+
}
|
|
3114
|
+
if (!isHttpUrl(params.video_url)) {
|
|
3115
|
+
if (!params.project_id) { console.error('Local video upload requires --project <id>.'); process.exit(1); }
|
|
3116
|
+
const valid = validateVideoFileForAnalysis(params.video_url);
|
|
3117
|
+
if (!valid.ok) { console.error(valid.error); process.exit(1); }
|
|
3118
|
+
params.video_url = await uploadFileViaSignedUrl(baseUrl, headers, params.project_id, params.video_url, valid.mime);
|
|
3119
|
+
if (!params.video_url) process.exit(1);
|
|
3120
|
+
}
|
|
3121
|
+
const receipt = await callMcpTool(baseUrl, headers, 'makaron_retake_video', params);
|
|
3122
|
+
const raw = receipt?.content?.find(c => c.type === 'text')?.text;
|
|
3123
|
+
let result;
|
|
3124
|
+
try { result = JSON.parse(raw); } catch { throw new Error(raw || 'Retake returned no receipt.'); }
|
|
3125
|
+
if (wait && result.success && result.taskId && !result.videoUrl) {
|
|
3126
|
+
const url = await pollVideo(baseUrl, headers, result.taskId, result.snapshotId);
|
|
3127
|
+
result = { ...result, status: url ? 'completed' : 'processing', videoUrl: url || undefined };
|
|
3128
|
+
}
|
|
3129
|
+
console.log(json ? JSON.stringify(result) : [result.message, result.taskId && `Task: ${result.taskId}`, result.videoUrl && `Video: ${result.videoUrl}`].filter(Boolean).join('\n'));
|
|
3130
|
+
if (!result.success) process.exitCode = 1;
|
|
3131
|
+
|
|
3088
3132
|
} else if (sub === 'create') {
|
|
3089
3133
|
let images = [];
|
|
3090
3134
|
const videos = [];
|
package/package.json
CHANGED
package/skills/makaron/SKILL.md
CHANGED
|
@@ -99,7 +99,7 @@ npx makaron-cli chat --project auto --image photo.jpg --json -b "make it cinemat
|
|
|
99
99
|
npx makaron-cli chat --project auto --image img1.jpg --image img2.jpg --json -b "combine these"
|
|
100
100
|
```
|
|
101
101
|
|
|
102
|
-
`chat` routes image and video models automatically. Use `--agent-model` only when the user explicitly asks to select or compare the reasoning/tool-calling Agent LLM. Accepted values are `auto`, the base model IDs (`gpt-6-luna`, `gpt-6-sol`, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-5.6-luna`, `grok-4.6`, `deepseek-v4-pro`, `deepseek-flash`), and the personal-plan routes (`gpt-6-luna-codex-subscription`, `gpt-6-sol-codex-subscription`, `gpt-5.6-terra-codex-subscription`, `gpt-5.6-sol-codex-subscription`, `gpt-5.6-luna-codex-subscription`). `auto` uses GPT-6 Luna through the Codex subscription for eligible accounts (including admins) and through Azure API otherwise; base GPT-5.6 and GPT-6 IDs select Azure API, while suffixed IDs explicitly request the personal plan where authorized. Never put an image or video model ID in `--agent-model`.
|
|
102
|
+
`chat` routes image and video models automatically. Use `--agent-model` only when the user explicitly asks to select or compare the reasoning/tool-calling Agent LLM. Accepted values are `auto`, the base model IDs (`gpt-6-luna`, `gpt-6.1-sol`, `gpt-6-sol`, `gpt-5.6-terra`, `gpt-5.6-sol`, `gpt-5.6-luna`, `grok-4.6`, `deepseek-v4-pro`, `deepseek-flash`), and the personal-plan routes (`gpt-6-luna-codex-subscription`, `gpt-6.1-sol-codex-subscription`, `gpt-6-sol-codex-subscription`, `gpt-5.6-terra-codex-subscription`, `gpt-5.6-sol-codex-subscription`, `gpt-5.6-luna-codex-subscription`). `auto` uses GPT-6 Luna through the Codex subscription for eligible accounts (including admins) and through Azure API otherwise; base GPT-5.6, GPT-6 and GPT-6.1 IDs select Azure API, while suffixed IDs explicitly request the personal plan where authorized. Never put an image or video model ID in `--agent-model`.
|
|
103
103
|
|
|
104
104
|
```bash
|
|
105
105
|
npx makaron-cli chat --project auto --agent-model deepseek-v4-pro --json -b "make a 20s badminton video"
|
|
@@ -341,8 +341,26 @@ npx makaron-cli video create --script "make it warmer and cinematic" --video htt
|
|
|
341
341
|
|
|
342
342
|
# 4. Check status
|
|
343
343
|
npx makaron-cli video status <taskId>
|
|
344
|
+
|
|
345
|
+
# Retake only a selected interval; publish the complete video to a project.
|
|
346
|
+
npx makaron-cli video retake --video input.mp4 --start 2 --end 4 --prompt "Make the book red" --model seedance-2.5 --project <id> --wait --json
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
Video segment editing also works directly from natural-language chat, without a GUI selection:
|
|
350
|
+
|
|
351
|
+
```bash
|
|
352
|
+
npx makaron-cli chat --project <id> "把 @1 的 18–21 秒切成多机位,其余不变"
|
|
353
|
+
npx makaron-cli chat --project <id> "把 @1 的最后三秒改成夜景"
|
|
344
354
|
```
|
|
345
355
|
|
|
356
|
+
Chat resolves the source/range, inspects the footage, expands the edit instruction, and delivers the complete video. It asks only when the source or scope remains ambiguous. Precise cuts, subtitles, dubbing and extension also start from chat. Local editing accepts a 0.1–15s range in a source up to 120s; longer scopes use the appropriate whole-video workflow rather than silently shortening the request.
|
|
357
|
+
|
|
358
|
+
Segment edits have two intents: modify changes the inside of the existing sequence and preserves its opening/closing states; replace discards a shot and does not lock its original endpoints. A camera change alone does not mean replacement. A same-request model comparison reuses the original source and interval, not the last generated result.
|
|
359
|
+
|
|
360
|
+
Chat segment inspection also reads source speech through ASR, reusing cached transcripts. It returns selected speech with source/output timecodes separately from adjacent speech, so the edit can follow the retained narration. Original audio is the default; Chat may select `audio_mode: generated` for requested new sound inside the selection. ASR does not guarantee music-beat or lip synchronization, and unavailable/untimed speech is reported rather than assigned invented timings. Direct `video retake` uses the caller's completed prompt and does not perform Agent interpretation or ASR planning.
|
|
361
|
+
|
|
362
|
+
`video retake` uses original-source seconds, accepts a 0.1–15s range in a source up to 120s, and automatically replaces only that range while preserving duration. Audio defaults to original; `--audio-mode generated` uses synchronized new audio inside the selection and source audio outside. Models: `fal-h3-max` (default), `seedance-2.5-eco` (preferred Seedance option), and explicitly selected native `seedance-2.5`. For sources shorter than 2 seconds, explicitly use Seedance. Local files require `--project`. Keep the root taskId and request ID; resume polling or replay the same request ID instead of paying for another generation.
|
|
363
|
+
|
|
346
364
|
`video create` returns a provider task id and does not create or update a Makaron project timeline. For project/timeline video editing, use:
|
|
347
365
|
|
|
348
366
|
```bash
|