aimakeall-mcp 0.11.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ Claude Code·Codex 같은 MCP 클라이언트에서 자연어로 AImakeAll 영
16
16
  "mcpServers": {
17
17
  "aimakeall": {
18
18
  "command": "npx",
19
- "args": ["-y", "aimakeall-mcp@0.11.0"],
19
+ "args": ["-y", "aimakeall-mcp@0.12.0"],
20
20
  "env": { "AIMAKEALL_PAT": "aio_pat_..." }
21
21
  }
22
22
  }
@@ -53,19 +53,29 @@ plan_shorts_video → (씬마다) generate_scene_image → get_generation_job(co
53
53
  - 원가 가시화: 생성 전 `estimate_video_cost`로 견적을 내고, 작업 후 `get_cost_report`로 실제 지출을 확인하세요. 어두운/검은 컷 의심 시 `verify_render_darkness`로 완성본 휘도를 실측할 수 있습니다.
54
54
  - 캐릭터 일관성: `plan_*`에 `characters`(이름·외모 앵커·의상 고정)를 넘기면 씬마다 identity lock 이 프롬프트에 강제 주입되고, `generate_scene_image`의 `identityLock`으로 재생성 시에도 유지됩니다.
55
55
 
56
- ## 제작 스타일 유지와 컷컴포저 편집 (0.11.0)
56
+ ## 제작 스타일 유지와 컷컴포저 편집 (0.12.0)
57
57
 
58
- 이 절의 도구는 MCP 0.11.0과 대응 서버·컴패니언 런타임이 필요합니다. 기존 MCP 설정의 버전을 `aimakeall-mcp@0.11.0`으로 바꾸고 클라이언트를 다시 연결하세요. 컴패니언에서 렌더 기능 업데이트를 요청하면 최신 설치 프로그램으로 업데이트하세요.
58
+ 아래의 이미지 소스·영상 표현 선택과 TTS 자동 우선순위는 MCP 0.12.0과 대응 서버·컴패니언 런타임이 필요합니다. 기존 MCP 설정의 버전을 `aimakeall-mcp@0.12.0`으로 바꾸고 클라이언트를 다시 연결하세요. 토큰은 그대로 사용할 수 있습니다. 컴패니언에서 렌더 기능 업데이트를 요청하면 최신 설치 프로그램으로 업데이트하세요.
59
59
 
60
60
  - `plan_shorts_video({categoryId:"viral-shorts",durationTargetSec:40,...})`는 `productionProfile`과 `productionPath`를 반환합니다. 이후 추천·TTS·스티치에 **가장 최근 반환된 `productionPath`**를 전달하세요. 제작별 불변 파일이라 동시에 만든 다른 영상의 스타일과 섞이지 않습니다.
61
61
  - 바이럴·커뮤니티 프로필은 `casual-social`을 사용합니다. 사건 소재라도 뉴스 브리핑형으로 자동 변경하지 않습니다. TTS의 `text`를 생략하면 저장된 기획 대본을 사용하며, 분명한 뉴스체 회귀는 과금 전 `422 NARRATION_STYLE_DRIFT`로 중단합니다. 자막은 실제 발화와 일치시켜야 하며 자극성을 위해 사실·인용·피해 내용을 꾸미면 안 됩니다.
62
- - 프로필이 있으면 보이스 추천 기본 공급자는 ElevenLabs입니다. 카탈로그에서 대화형/소셜 특성을 우선하되 실제 청취 결과라고 주장하지 않습니다. `speed`, `stability`, `style`, `similarityBoost`, `speakerBoost`, `modelId`는 명시적으로 조절할 수 있습니다.
62
+ - 보이스 공급자 자동 선택은 Typecast → ElevenLabs → Supertone → Edge TTS → Supertonic2 순서입니다. 키/실행 가능 상태를 먼저 확인하며 유료 합성 오류 뒤 다른 공급자로 몰래 재합성하지 않습니다. 대화형/소셜 목소리를 우선하되 실제 청취 결과라고 주장하지 않습니다. ElevenLabs 외의 공급자 자막은 실제 오디오 길이 기반 추정 타이밍임을 표시하므로 최종 동기화를 확인하세요.
63
63
  - `list_composer_presets`에서 실제 편집기의 프리셋과 모션을 조회하세요. `titleStylePresetId`, `subtitleStylePresetId`, `subtitleLines[].stylePresetId`, `titleStyle`, `subtitleStyle`, `subtitleLines[].style`로 서로 다른 제목/본문 계층과 장면별 강조를 설정합니다. 명시 스타일 > 선택 프리셋 > 프로필 기본값 순서입니다.
64
64
  - `search_composer_assets` 또는 무료 `recommend_composer_overlays`로 관련 소재를 검토하고 선택한 `assetId`에 `startSec/endSec`, `xPct/yPct`, `widthPct`를 붙여 `stitch_timeline.overlays`로 전달하세요. 좌표는 화면의 백분율 중심점이며 `opacity`는 0~100입니다. SVG 아이콘·이모지, GIF/비디오, 효과음을 타임라인 레이어로 처리합니다. 사건·피해자에 조롱 밈을 강제로 붙이지 않으며, 출처·라이선스 확인이 필요한 소재는 공개 전에 검토해야 합니다.
65
65
  - 목표 길이는 실제 TTS 길이로 덮어쓰지 않습니다. 허용 오차 `max(2초, 목표의 10%)`를 벗어나면 경고하고 렌더 전에 차단합니다. 자동 유료 재합성을 하지 않으며, 무음이나 억지 영상 늘이기로 통과시키지 마세요.
66
66
  - 커뮤니티/바이럴 제작은 라이브러리 검토 없이 기본 자막만으로 끝내지 않습니다. 오버레이가 없으면 `decorationRationale`에 적합한 소재를 생략한 이유를 기록해야 렌더할 수 있습니다. 관련 없는 장식이나 부적절한 밈을 억지로 넣는 것은 해결이 아닙니다.
67
67
  - `review_production`은 구조만 점검합니다. `render_result`의 `qualityStatus`가 `needs-visual-and-listening-review`이면 아직 시각·청취 검증 전입니다. 최종 영상에서 강조색/글자 크기/모션/소재 겹침과 목소리를 직접 확인한 뒤에만 완성 품질을 보고하세요.
68
68
 
69
+ ## 장면 이미지 준비와 영상 표현 선택
70
+
71
+ 다음 설정은 0.12.0부터 지원합니다. 기존 0.11.0 클라이언트는 위의 버전 고정 설정을 바꾸고 다시 연결해야 새 입력과 도구 안내를 사용할 수 있습니다. 이미지 모션 렌더에는 해당 기능을 지원하는 최신 컴패니언 런타임도 필요합니다.
72
+
73
+ - `plan_shorts_video`, `plan_commerce_video`, `plan_music_video`에서 `imageSourceMode: "web" | "ai" | "mixed"`와 `visualMotionMode: "image-motion" | "i2v"`를 독립적으로 선택합니다. 생략 시 기존 `ai + i2v`입니다. 반환된 `productionPath`를 이미지·영상·TTS·스티치에 계속 전달하면 다른 작업과 선택이 섞이지 않습니다.
74
+ - 웹은 사용자 계정의 Pexels/Pixabay 키로 장면별 `webSearchQuery`를 검색합니다. `web`은 결과가 없으면 중단하고, `mixed`만 결과가 없을 때 AI로 이어집니다. 검색 요청 자체가 실패하면 먼저 오류를 확인합니다. `ai`는 웹 검색을 실행하지 않습니다. 여기서 ‘AI만’은 주 장면 이미지 기준이며 편집용 아이콘·이모지 등은 별도입니다.
75
+ - 웹 결과는 즉시 `imageUrl`, `sceneVisual`, 출처·작가·이용 조건을 반환합니다. AI 생성만 `generationJobId`를 조회해야 합니다. `excludeImageUrls` 또는 직전 이미지 준비 결과의 `productionPath`로 웹 이미지 반복을 피하세요. 웹 사진이 실제 사건 사진이라는 뜻은 아니며, 최종 공개 전에 내용과 출처·사용 조건을 검토하세요.
76
+ - `image-motion`에서는 `generate_scene_video_prompt`와 `generate_scene_video`를 실행하지 않습니다. `stitch_timeline({productionPath,sceneVisuals:[{kind:"image",idx:0,url:imageUrl,durationSec:6,source:"web",sourcePage,attribution,license}],autoCompose:true})`처럼 이미지를 바로 넘깁니다. `list_composer_presets.imageMotions`의 ID와 0~100 강도를 `imageMotionPreset`/`imageMotionIntensity`로 지정할 수 있습니다. 기존 `sceneVideos` 입력도 유지합니다.
77
+ - `autoCompose:true`는 실제 인스펙터 모션·장면별 자막 계층·문맥에 맞는 안전한 라이브러리 소재를 자동 배치합니다. 명시한 스타일·오버레이가 우선이고, 필요하면 `false`로 수동 편집을 유지합니다. `compositionReport.autoComposition`에서 적용·생략 이유를 확인하세요. 모든 장면에 밈이나 효과음을 강제로 넣지 않으며 권리가 확인되지 않은 자산은 자동 삽입하지 않습니다.
78
+
69
79
  ## 채널 규격 지문 (벤치마킹)
70
80
 
71
81
  ```
@@ -27,8 +27,10 @@ import { fetchGenerationJob, GENERATION_JOB_PATH, submitGenerationJob } from "./
27
27
  import { loudnessDbfsFromWav, silenceRatioFromWav } from "./wav-dsp.mjs";
28
28
  import { trimWavToSeconds } from "./wav-trim.mjs";
29
29
  import { registerComposerTools } from "./composer-tools.mjs";
30
- import { productionFields, composerTextStyle, composerOverlay } from "./production-schemas.mjs";
30
+ import { productionFields, visualProductionFields, sceneVisual, composerTextStyle, composerOverlay } from "./production-schemas.mjs";
31
31
  import { saveProductionContext, readProductionContext, durationReview, reviewProductionPayload } from "./production-context.mjs";
32
+ import { TTS_PROVIDERS, buildTtsRequestBody, estimateTtsSubtitleLines, requestTtsBinary, resolveTtsSelection } from "./tts-workflow.mjs";
33
+ import { assertI2vAllowed, findSceneWebImage, resolveVisualSettings, sceneWebSearchQuery } from "./visual-production.mjs";
32
34
 
33
35
  // 여러 파일을 base64 로 싣는 요청의 인코딩 크기 합이 서버 JSON 상한을 넘지 않게 사전 검사.
34
36
  // 서버가 전체 업로드를 받은 뒤 413 을 내는 것을 막고 행동 가능한 한국어 안내를 준다.
@@ -165,6 +167,7 @@ function trimPlanScenes(scenes) {
165
167
  durationSec: scene?.durationSec,
166
168
  idx: scene?.idx,
167
169
  imagePrompt: scene?.imagePrompt,
170
+ webSearchQuery: sceneWebSearchQuery(scene),
168
171
  label: scene?.label,
169
172
  ratio: scene?.ratio,
170
173
  sceneNarration: scene?.sceneNarration,
@@ -189,6 +192,7 @@ export function registerCloudTools(server, config, api) {
189
192
  const resultBody = { generationJobId: job.id, status: job.status, kind: job.kind,
190
193
  model: payload.model, taskId: payload.taskId, providerTasks: job.providerTasks };
191
194
  if (job.kind === "image") {
195
+ resultBody.source = "ai";
192
196
  resultBody.imageUrl = payload.imageUrl || undefined;
193
197
  let inlineImage = null;
194
198
  if (!payload.imageUrl && payload.imageDataUrl) {
@@ -285,10 +289,11 @@ export function registerCloudTools(server, config, api) {
285
289
  videoModel: z.string().optional().describe("기본 kie-grok-imagine"),
286
290
  videoSecondsPerScene: z.number().min(0).max(30).optional().describe("씬당 영상 초 (0이면 영상 미포함)"),
287
291
  videoQuality: z.string().optional().describe("기본 720p"),
288
- ttsProvider: z.enum(["typecast", "elevenlabs", "supertone"]).optional().describe("기본 typecast"),
292
+ ttsProvider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("기본 계정별 자동 우선순위. Edge/Supertonic2는 TTS 공급자 비용 0"),
289
293
  ttsChars: z.number().int().min(0).max(20000).optional().describe("내레이션 글자 수"),
290
294
  },
291
295
  wrapCloudHandler(config, async (input) => {
296
+ if (input.ttsChars > 0) input = { ...input, ttsProvider: (await resolveTtsSelection(api, config, { provider: input.ttsProvider })).provider };
292
297
  const payload = await api.request("/api/tracker/cost/estimate", { body: input, method: "POST", timeoutMs: 30_000 });
293
298
  return jsonResult(payload);
294
299
  }),
@@ -299,6 +304,7 @@ export function registerCloudTools(server, config, api) {
299
304
  "쇼츠 영상 기획(시나리오·씬별 이미지 프롬프트·내레이션)을 생성합니다. topic 또는 copy 중 하나는 필수. 결과 scenes의 imagePrompt는 generate_scene_image로, skill은 씬 영상 프롬프트의 videoStylePreset으로, emphasisKeywords는 stitch_timeline의 emphasisKeywords로 이어집니다(하단자막 단어 강조).",
300
305
  {
301
306
  productionProfile: productionFields.productionProfile,
307
+ ...visualProductionFields,
302
308
  durationTargetSec: z.number().positive().max(180).optional().describe("사용자가 요청한 목표 길이(예:40). 음성 길이로 덮어쓰지 않습니다."),
303
309
  topic: z.string().optional().describe("영상 주제"),
304
310
  copy: z.string().optional().describe("핵심 카피/대사"),
@@ -328,12 +334,14 @@ export function registerCloudTools(server, config, api) {
328
334
  forbidden: z.string().optional().describe("금지 변형 (기본: different face, different hairstyle)"),
329
335
  })).max(4).optional().describe("등장 캐릭터 — 서버가 씬마다 identity_lock 을 imagePrompt 에 강제 주입해 캐릭터 일관성을 잠급니다"),
330
336
  },
331
- wrapCloudHandler(config, async ({ topic = "", copy = "", categoryId = "community-shorts", targetCustomer = "", tone = "자동 추천", sceneCount = 0, characters = [], channelFingerprint = null, productionProfile, durationTargetSec }) => {
337
+ wrapCloudHandler(config, async ({ topic = "", copy = "", categoryId = "community-shorts", targetCustomer = "", tone = "자동 추천", sceneCount = 0, characters = [], channelFingerprint = null, productionProfile, durationTargetSec, imageSourceMode, visualMotionMode }) => {
332
338
  if (!String(topic).trim() && !String(copy).trim()) {
333
339
  return textResult("topic 또는 copy 중 하나는 입력해야 합니다.", { isError: true });
334
340
  }
341
+ const visualSettings = resolveVisualSettings(null, { productionProfile, imageSourceMode, visualMotionMode });
335
342
  const payload = await api.request("/api/tracker/shorts-video/plan", {
336
343
  body: {
344
+ ...visualSettings,
337
345
  categoryId,
338
346
  productionProfile,
339
347
  durationTargetSec,
@@ -352,15 +360,17 @@ export function registerCloudTools(server, config, api) {
352
360
  timeoutMs: 240_000,
353
361
  });
354
362
  const productionPath = saveProductionContext(config.stateDir, {
355
- topic, productionProfile: payload?.productionProfile, narrationScript: payload?.narrationScript,
363
+ topic, visualSettings, autoCompose: productionProfile?.autoComposeEnabled ?? true,
364
+ productionProfile: payload?.productionProfile, narrationScript: payload?.narrationScript,
356
365
  narrationSource: payload?.narrationSource, tagline: payload?.tagline, emphasisKeywords: payload?.emphasisKeywords || [],
357
366
  channelFingerprint, scenes: trimPlanScenes(payload?.scenes),
358
367
  });
359
368
  return jsonResult({
360
369
  productionPath,
370
+ ...visualSettings,
361
371
  productionProfile: payload?.productionProfile,
362
372
  narrationSource: payload?.narrationSource,
363
- next: "productionPath를 recommend_voice → tts_narration_with_captions → stitch_timeline에 전달하세요. 대본을 뉴스체로 바꾸지 말고, list_composer_presets/search_composer_assets로 편집 소재도 선택하세요.",
373
+ next: `productionPath를 이미지 준비·보이스 추천·TTS·스티치에 계속 전달하세요. ${visualSettings.visualMotionMode === "image-motion" ? "I2V 프롬프트/영상 생성은 생략하고 이미지 URL을 stitch_timeline.sceneVisuals(kind=image)에 넣으세요." : "이미지 확인 후 I2V 영상 생성으로 진행하세요."} 대본을 뉴스체로 바꾸지 말고 컷컴포저 소재도 활용하세요.`,
364
374
  charactersUsed: Array.isArray(payload?.charactersUsed) ? payload.charactersUsed : undefined,
365
375
  ctaText: payload?.ctaText,
366
376
  emphasisKeywords: Array.isArray(payload?.emphasisKeywords) ? payload.emphasisKeywords : [],
@@ -455,6 +465,8 @@ export function registerCloudTools(server, config, api) {
455
465
  "plan_commerce_video",
456
466
  "제품 홍보 영상 기획을 생성합니다. 제품 사진 파일 경로가 최소 1장 필요합니다 (이 PC의 로컬 경로).",
457
467
  {
468
+ ...visualProductionFields,
469
+ productionProfile: productionFields.productionProfile,
458
470
  productName: z.string().describe("제품명 (필수)"),
459
471
  productImagePaths: z.array(z.string()).min(1).describe("제품/자산 사진의 로컬 파일 경로 (1장 이상)"),
460
472
  modelImagePath: z.string().optional().describe("모델 사진 로컬 경로 (선택)"),
@@ -472,7 +484,8 @@ export function registerCloudTools(server, config, api) {
472
484
  forbidden: z.string().optional().describe("금지 변형 (기본: different face, different hairstyle)"),
473
485
  })).max(4).optional().describe("등장 캐릭터 — 서버가 씬마다 identity_lock 을 imagePrompt 에 강제 주입해 캐릭터 일관성을 잠급니다"),
474
486
  },
475
- wrapCloudHandler(config, async ({ productName, productImagePaths, modelImagePath = "", description = "", targetCustomer = "", tone = "자동 추천", categoryId = "ecommerce", sceneCount = 6, researchBrief = "" , characters = [] }) => {
487
+ wrapCloudHandler(config, async ({ productName, productImagePaths, modelImagePath = "", description = "", targetCustomer = "", tone = "자동 추천", categoryId = "ecommerce", sceneCount = 6, researchBrief = "" , characters = [], productionProfile, imageSourceMode, visualMotionMode }) => {
488
+ const visualSettings = resolveVisualSettings(null, { productionProfile, imageSourceMode, visualMotionMode });
476
489
  // 확장자 검증(임의 파일 업로드 차단) + 합산 크기 예산(서버 413 사전 차단).
477
490
  for (const filePath of productImagePaths) assertAllowedInputFile(filePath, ALLOWED_IMAGE_EXTS, { kind: "이미지" });
478
491
  if (modelImagePath) assertAllowedInputFile(modelImagePath, ALLOWED_IMAGE_EXTS, { kind: "모델 이미지" });
@@ -480,6 +493,7 @@ export function registerCloudTools(server, config, api) {
480
493
 
481
494
  const payload = await api.request("/api/tracker/commerce-video/plan", {
482
495
  body: {
496
+ ...visualSettings, productionProfile,
483
497
  categoryId,
484
498
  description,
485
499
  modelImage: modelImagePath ? fileToImagePayload(modelImagePath) : null,
@@ -497,7 +511,14 @@ export function registerCloudTools(server, config, api) {
497
511
  method: "POST",
498
512
  timeoutMs: 300_000,
499
513
  });
514
+ const productionPath = saveProductionContext(config.stateDir, {
515
+ topic: productName, visualSettings, autoCompose: productionProfile?.autoComposeEnabled ?? true,
516
+ productionProfile: payload?.productionProfile || productionProfile,
517
+ tagline: payload?.tagline, narrationScript: payload?.narrationScript,
518
+ scenes: trimPlanScenes(payload?.scenes), featureKey: "commerceVideo",
519
+ });
500
520
  return jsonResult({
521
+ productionPath, ...visualSettings, productionProfile: payload?.productionProfile || productionProfile,
501
522
  analysis: trimAnalysis(payload?.analysis),
502
523
  ctaText: payload?.ctaText,
503
524
  ok: payload?.ok,
@@ -510,8 +531,13 @@ export function registerCloudTools(server, config, api) {
510
531
 
511
532
  server.tool(
512
533
  "generate_scene_image",
513
- "씬 이미지 생성 작업을 접수하고 즉시 generationJobId를 반환합니다. get_generation_job으로 completed 결과의 imageUrl을 받은 뒤 씬 영상 입력에 쓰세요. 첫 씬/캐릭터/제품 이미지를 referenceImageUrls·referenceImagePaths로 전달하면 일관성 유지에 도움이 됩니다. 완료 이미지는 직접 확인하거나 verify_scene_image로 검증하세요.",
534
+ "선택한 방식으로 장면 이미지를 준비합니다. productionPath의 web은 사용자 키로 Pexels/Pixabay 검색만, ai는 AI 생성만, mixed는 웹 검색 후 없을 때 AI 생성합니다. 웹 결과는 즉시 imageUrl·출처를 반환하며 AI만 generationJobId를 get_generation_job으로 조회합니다. image-motion이면 I2V를 생략하고 sceneVisual을 스티치에 사용하세요.",
514
535
  {
536
+ ...productionFields,
537
+ ...visualProductionFields,
538
+ sceneIdx: z.number().int().min(0).optional().describe("기획 장면 idx. 저장된 장면 검색어와 출처를 연결합니다."),
539
+ webSearchQuery: z.string().max(160).optional().describe("장면에 맞는 짧은 웹 검색 키워드. plan.scenes[].webSearchQuery를 사용하세요."),
540
+ excludeImageUrls: z.array(z.string()).max(100).optional().describe("이미 사용한 웹 이미지 URL. 반복 장면을 피합니다."),
515
541
  requestId: z.string().optional().describe("접수 응답 유실 후 같은 작업을 재확인할 때 기존 generationJobId를 지정합니다. 동일 ID·입력은 중복 생성하지 않습니다."),
516
542
  prompt: z.string().describe("이미지 프롬프트 (plan 결과의 imagePrompt)"),
517
543
  aspectRatio: z.string().optional().describe("기본 9:16"),
@@ -522,7 +548,46 @@ export function registerCloudTools(server, config, api) {
522
548
  referenceImageUrls: z.array(z.string()).max(6).optional().describe("참조 이미지 URL (씬1 앵커 등)"),
523
549
  referenceImagePaths: z.array(z.string()).max(4).optional().describe("참조 이미지 로컬 경로 (제품 사진 등)"),
524
550
  },
525
- wrapCloudHandler(config, async ({ requestId, prompt, aspectRatio = "9:16", model = "gpt-image-2-beta", resolution = "1K", referenceImageUrls = [], referenceImagePaths = [], returnImage = true, identityLock = "" }) => {
551
+ wrapCloudHandler(config, async ({ requestId, prompt, aspectRatio = "9:16", model = "gpt-image-2-beta", resolution = "1K", referenceImageUrls = [], referenceImagePaths = [], returnImage = true, identityLock = "", productionPath, productionProfile, imageSourceMode, visualMotionMode, sceneIdx = 0, webSearchQuery, excludeImageUrls = [] }) => {
552
+ const context = readProductionContext(config.stateDir, productionPath);
553
+ const visualSettings = resolveVisualSettings(context, { productionProfile, imageSourceMode, visualMotionMode });
554
+ productionProfile = context?.productionProfile || productionProfile;
555
+ const plannedScene = context?.scenes?.find((entry) => entry.idx === sceneIdx) || {};
556
+ const saveSource = (source) => saveProductionContext(config.stateDir, {
557
+ ...context, productionProfile, visualSettings,
558
+ sceneSources: [...(context?.sceneSources || []).filter((entry) => entry.idx !== sceneIdx), { ...source, idx: sceneIdx }],
559
+ });
560
+ let fallbackReason;
561
+ if (visualSettings.imageSourceMode !== "ai") {
562
+ // A retry of a possibly accepted paid job must recover it before another
563
+ // stock selection can hide it. Unknown IDs continue the original flow.
564
+ if (requestId) {
565
+ try {
566
+ const job = await fetchGenerationJob(api, requestId);
567
+ if (job.kind !== "image") throw new Error("requestId는 이미지 생성 작업 ID여야 합니다. 영상 작업은 get_generation_job으로 조회하세요.");
568
+ return job.status === "completed" ? completedGenerationResult(job, returnImage)
569
+ : jsonResult({ generationJobId: job.id, status: job.status, kind: job.kind, existing: true, nextTool: "get_generation_job" });
570
+ } catch (error) { if (Number(error?.status) !== 404) throw error; }
571
+ }
572
+ const query = webSearchQuery || sceneWebSearchQuery(plannedScene) || sceneWebSearchQuery({ imagePrompt: prompt });
573
+ const selected = await findSceneWebImage(api, {
574
+ query,
575
+ excludeImageUrls: [...excludeImageUrls, ...(context?.sceneSources || []).filter((entry) => entry.idx !== sceneIdx).map((entry) => entry.url).filter(Boolean)],
576
+ });
577
+ if (selected) {
578
+ const visual = { ...selected, idx: sceneIdx, label: plannedScene.label || selected.label, durationSec: Number(plannedScene.durationSec) || 6 };
579
+ return jsonResult({ ok: true, status: "completed", ...visualSettings, source: "web", imageUrl: selected.url,
580
+ sceneVisual: visual, sourcePage: selected.sourcePage, attribution: selected.attribution, license: selected.license,
581
+ productionPath: saveSource(visual),
582
+ guidance: visualSettings.visualMotionMode === "image-motion" ? "sceneVisual을 stitch_timeline.sceneVisuals에 넣으세요. I2V 호출은 하지 않습니다." : "웹 이미지의 내용·출처를 확인한 뒤 I2V 입력으로 사용하세요. 생성 영상에도 이 출처 정보를 보존하세요.",
583
+ });
584
+ }
585
+ if (visualSettings.imageSourceMode === "web") {
586
+ return { ...jsonResult({ ok: false, code: "WEB_IMAGE_NOT_FOUND", ...visualSettings,
587
+ message: "사용 가능한 새 웹 이미지를 찾지 못했습니다. 스톡 검색 키·검색어를 확인하거나 직접 이미지를 선택하세요. 웹만 선택했으므로 AI 생성은 실행하지 않았습니다." }), isError: true };
588
+ }
589
+ fallbackReason = "혼합 모드: 사용 가능한 새 웹 이미지가 없어 선택한 AI 생성으로 진행합니다.";
590
+ }
526
591
  // identity_lock 리터럴 부착 — 이미 포함돼 있으면 재부착하지 않는다(멱등: 프롬프트 캐시 보존).
527
592
  const lock = String(identityLock || "").trim();
528
593
  if (lock && !prompt.includes(lock)) prompt = `${prompt.trim()}, ${lock}`;
@@ -536,18 +601,22 @@ export function registerCloudTools(server, config, api) {
536
601
  ...referenceImageUrls.map((url, index) => ({ name: `ref-${index + 1}`, previewUrl: url })),
537
602
  ].slice(0, 6);
538
603
 
539
- return jsonResult(await submitGenerationJob(api, {
604
+ const job = await submitGenerationJob(api, {
540
605
  kind: "image", includePreview: returnImage, requestId,
541
606
  input: {
607
+ productionProfile, ...visualSettings,
542
608
  aspectRatio,
543
609
  model,
544
610
  prompt,
545
611
  referenceImages,
546
- referenceSearchEnabled: true,
612
+ referenceSearchEnabled: false,
547
613
  resolution,
548
- webSearchEnabled: true,
614
+ webSearchEnabled: false,
549
615
  },
550
- }));
616
+ });
617
+ return jsonResult({ ...job, ...visualSettings, source: "ai", fallbackReason,
618
+ productionPath: saveSource({ source: "ai", generationJobId: job.generationJobId }),
619
+ });
551
620
  }),
552
621
  );
553
622
 
@@ -555,6 +624,8 @@ export function registerCloudTools(server, config, api) {
555
624
  "generate_scene_video_prompt",
556
625
  "씬 이미지 기반 영상 프롬프트를 생성합니다. 결과 videoPrompt를 generate_scene_video의 prompt로 넘기세요.",
557
626
  {
627
+ ...productionFields,
628
+ ...visualProductionFields,
558
629
  imagePrompt: z.string().describe("씬의 이미지 프롬프트"),
559
630
  sceneImageUrl: z.string().optional().describe("generate_scene_image가 반환한 imageUrl"),
560
631
  sceneScript: z.string().optional().describe("씬 대사/설명 (plan 결과의 label)"),
@@ -565,9 +636,13 @@ export function registerCloudTools(server, config, api) {
565
636
  kind: z.enum(["shorts_scene", "commerce_video_scene", "music_video_scene"]).optional(),
566
637
  modelId: z.string().optional().describe("대상 영상 모델 (기본 kie-grok-imagine — generate_scene_video와 동일하게 맞추세요)"),
567
638
  },
568
- wrapCloudHandler(config, async ({ imagePrompt, sceneImageUrl = "", sceneScript = "", sceneTitle = "", aspectRatio = "9:16", durationSec = 8, videoStylePreset = "", kind = "shorts_scene", modelId = "kie-grok-imagine" }) => {
639
+ wrapCloudHandler(config, async ({ imagePrompt, sceneImageUrl = "", sceneScript = "", sceneTitle = "", aspectRatio = "9:16", durationSec = 8, videoStylePreset = "", kind = "shorts_scene", modelId = "kie-grok-imagine", productionPath, productionProfile, imageSourceMode, visualMotionMode }) => {
640
+ const context = readProductionContext(config.stateDir, productionPath);
641
+ const visualSettings = resolveVisualSettings(context, { productionProfile, imageSourceMode, visualMotionMode });
642
+ assertI2vAllowed(visualSettings);
569
643
  const payload = await api.request("/api/tracker/gemini/storyboard-video-prompt", {
570
644
  body: {
645
+ productionProfile: context?.productionProfile || productionProfile, ...visualSettings,
571
646
  aspectRatio,
572
647
  durationLabel: `${Math.max(1, Math.round(durationSec))}s`,
573
648
  generationPlan: {
@@ -596,6 +671,8 @@ export function registerCloudTools(server, config, api) {
596
671
  "generate_scene_video",
597
672
  "씬 영상 생성 작업을 접수하고 즉시 generationJobId를 반환합니다. get_generation_job으로 completed가 될 때까지 조회한 뒤 videoUrl을 stitch_timeline에 넣으세요. 접수만 된 상태에서 새로 생성하지 마세요.",
598
673
  {
674
+ ...productionFields,
675
+ ...visualProductionFields,
599
676
  requestId: z.string().optional().describe("접수 응답 유실 후 동일 입력을 재접수할 때 기존 generationJobId를 지정합니다."),
600
677
  prompt: z.string().describe("영상 프롬프트 (generate_scene_video_prompt의 videoPrompt)"),
601
678
  sceneImageUrl: z.string().optional().describe("씬 이미지 URL (i2v 입력)"),
@@ -605,10 +682,14 @@ export function registerCloudTools(server, config, api) {
605
682
  quality: z.string().optional().describe("기본 720p"),
606
683
  returnFrames: z.boolean().optional().describe("기본 true — 결과 영상의 대표 프레임 3장을 직접 보고 QC 할 수 있게 이미지 블록으로 반환"),
607
684
  },
608
- wrapCloudHandler(config, async ({ requestId, prompt, sceneImageUrl = "", aspectRatio = "9:16", durationSec = 8, modelId = "kie-grok-imagine", quality = "720p", returnFrames = true }) => {
685
+ wrapCloudHandler(config, async ({ requestId, prompt, sceneImageUrl = "", aspectRatio = "9:16", durationSec = 8, modelId = "kie-grok-imagine", quality = "720p", returnFrames = true, productionPath, productionProfile, imageSourceMode, visualMotionMode }) => {
686
+ const context = readProductionContext(config.stateDir, productionPath);
687
+ const visualSettings = resolveVisualSettings(context, { productionProfile, imageSourceMode, visualMotionMode });
688
+ assertI2vAllowed(visualSettings);
609
689
  return jsonResult(await submitGenerationJob(api, {
610
690
  kind: "video", includePreview: returnFrames, requestId,
611
691
  input: {
692
+ productionProfile: context?.productionProfile || productionProfile, ...visualSettings,
612
693
  aspectRatio,
613
694
  durationLabel: `${Math.max(1, Math.ceil(durationSec))}s`,
614
695
  generationPlan: {
@@ -630,11 +711,11 @@ export function registerCloudTools(server, config, api) {
630
711
 
631
712
  server.tool(
632
713
  "tts_narration",
633
- "내레이션 TTS를 합성해 이 PC에 mp3로 저장하고 파일 경로를 반환합니다. stitch_timeline의 ttsAudioPath로 쓰세요. 사용자가 보이스를 직접 지정하지 않았다면 먼저 recommend_voice로 주제·샘플 영상에 맞는 보이스를 자동 선정해 voice에 넣으세요.",
714
+ "내레이션 TTS를 합성해 이 PC에 오디오 파일로 저장합니다. 자동 순서: Typecast → ElevenLabs → Supertone → Edge TTS → Supertonic2. 키/런타임이 없을 때 다음 공급자를 고려하며, 합성 실패 후 다른 유료 공급자로 재시도하지 않습니다. 먼저 recommend_voice로 제작 톤에 맞는 보이스를 선택하세요.",
634
715
  {
635
716
  ...productionFields,
636
717
  text: z.string().optional().describe("생략 시 productionPath의 기획 대본"),
637
- provider: z.enum(["typecast", "elevenlabs"]).optional().describe("기본 typecast"),
718
+ provider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("기본 자동: Typecast→ElevenLabs→Supertone→Edge→Supertonic2. 명시 선택/저장된 보이스는 유지"),
638
719
  voice: z.string().optional().describe("보이스 이름(라벨)"),
639
720
  speed: z.number().optional().describe("기본 1"),
640
721
  stability: z.number().min(0).max(1).optional(), style: z.number().min(0).max(1).optional(),
@@ -644,57 +725,27 @@ export function registerCloudTools(server, config, api) {
644
725
  const context = readProductionContext(config.stateDir, productionPath);
645
726
  productionProfile = context?.productionProfile || productionProfile;
646
727
  text ??= context?.narrationScript;
647
- provider ||= context?.voice?.provider || (productionProfile ? "elevenlabs" : "typecast");
728
+ const selection = await resolveTtsSelection(api, config, { provider, context });
729
+ provider = selection.provider;
648
730
  if (!voice && context?.voice?.provider && context.voice.provider !== provider) return textResult("보이스 공급자가 다릅니다. 선택한 provider용 voice를 다시 선정하세요.", { isError: true });
649
731
  voice ||= context?.voice?.voiceId || "";
650
732
  if (!text?.trim()) return textResult("대본(text/productionPath)이 필요합니다.", { isError: true });
651
733
  if (productionProfile?.deliveryStyle === "casual-social" && !voice) return textResult("먼저 recommend_voice에 productionPath를 전달해 커뮤니티형 보이스를 선정하세요.", { isError: true });
652
- const body = provider === "typecast"
653
- ? {
654
- productionProfile,
655
- breath: 0.3,
656
- costEventId: createUsageEventId("mcp-tts-typecast"),
657
- emotion: "neutral",
658
- fileName: "mcp-narration",
659
- language: "kor",
660
- modelId: "",
661
- pitch: 0,
662
- smartEmotion: true,
663
- speed,
664
- text,
665
- voice,
666
- volume: 100,
667
- }
668
- : {
669
- productionProfile,
670
- costEventId: createUsageEventId("mcp-tts-elevenlabs"),
671
- fileName: "mcp-narration",
672
- languageCode: "ko",
673
- modelId,
674
- similarityBoost,
675
- speakerBoost,
676
- speed,
677
- stability,
678
- style,
679
- text,
680
- voice,
681
- };
682
- const result = await api.requestBinary(`/api/tracker/tts/${provider}`, {
683
- body,
684
- method: "POST",
685
- timeoutMs: 120_000,
686
- });
687
- const saved = saveMediaBuffer(config.stateDir, "narration.mp3", result.buffer, { contentType: result.contentType });
734
+ const body = buildTtsRequestBody(provider, { productionProfile, costEventId: createUsageEventId(`mcp-tts-${provider}`), modelId, similarityBoost, speakerBoost, speed, stability, style, text, voice });
735
+ const result = await requestTtsBinary(api, config, selection, body);
736
+ const saved = saveMediaBuffer(config.stateDir, provider === "supertonic2" ? "narration.wav" : "narration.mp3", result.buffer, { contentType: result.contentType });
688
737
  const durationSec = audioDurationSecFromFile(saved.filePath);
689
738
  return jsonResult({
690
- productionPath: saveProductionContext(config.stateDir, { ...context, productionProfile, narrationScript: text, tts: { filePath: saved.filePath, durationSec, voiceId: voice } }),
739
+ productionPath: saveProductionContext(config.stateDir, { ...context, productionProfile, narrationScript: text, tts: { provider, filePath: saved.filePath, durationSec, voiceId: voice } }),
740
+ provider,
741
+ executionTarget: selection.executionTarget,
691
742
  durationWarning: durationReview(productionProfile?.targetDurationSec, durationSec),
692
743
  bytes: saved.bytes,
693
744
  // 실측 mp3 길이 — stitch_timeline 의 audioDurationSec 로 그대로 쓰면 타임라인이 정확하다.
694
745
  durationSec,
695
746
  filePath: saved.filePath,
696
747
  voiceName: decodeHeaderValue(
697
- result.headers.get("x-typecast-voice-name") || result.headers.get("x-elevenlabs-voice-name") || "",
748
+ result.headers.get("x-typecast-voice-name") || result.headers.get("x-elevenlabs-voice-name") || result.headers.get("x-supertone-voice-name") || result.headers.get("x-edge-tts-voice-name") || result.headers.get("x-supertonic2-voice-name") || "",
698
749
  ),
699
750
  });
700
751
  }),
@@ -768,9 +819,12 @@ export function registerCloudTools(server, config, api) {
768
819
 
769
820
  server.tool(
770
821
  "stitch_timeline",
771
- "씬 영상들과 오디오를 서버에서 타임라인 매니페스트로 합칩니다. 결과는 파일 핸들(payloadPath)로 반환되며, 이를 render_start에 넘기면 이 PC에서 mp4가 렌더됩니다. 서버는 렌더하지 않습니다. 하단자막 강조: emphasisKeywords 를 주면 자막 라인 안의 해당 단어만 다른 색/크기로 강조되고, subtitleLines[].style 로 라인별 디자인도 바꿀 수 있습니다.",
822
+ "씬 이미지/영상과 오디오를 타임라인 매니페스트로 합칩니다. image-motion은 sceneVisuals(kind=image)에 인스펙터 imageMotionPreset/Intensity를 적용해 I2V 없이 렌더합니다. autoCompose 기본 true는 문맥별 모션·자막 디자인·적합한 라이브러리 소재를 배치하며 명시 편집을 우선합니다. 결과 payloadPath를 로컬 render_start에 넘기세요.",
772
823
  {
773
824
  ...productionFields,
825
+ ...visualProductionFields,
826
+ autoCompose: z.boolean().optional().describe("기본 true. 실제 인스펙터 모션·자막 계층·문맥별 라이브러리 자동 편집. false는 수동 편집을 유지합니다."),
827
+ sceneVisuals: z.array(sceneVisual).min(1).max(200).optional().describe("장면 이미지/영상과 출처 메타데이터. sceneVideos와 동시에 전달하지 마세요."),
774
828
  titleStylePresetId: z.string().optional(), subtitleStylePresetId: z.string().optional(),
775
829
  titleStyle: composerTextStyle.optional(), subtitleStyle: composerTextStyle.optional(),
776
830
  titleSegments: z.array(z.object({ text: z.string(), color: z.string().optional() })).max(20).optional(),
@@ -781,7 +835,7 @@ export function registerCloudTools(server, config, api) {
781
835
  idx: z.number().int(),
782
836
  label: z.string().optional(),
783
837
  url: z.string(),
784
- })).min(1).describe("씬 순서대로 idx=0부터. url은 generate_scene_video의 videoUrl"),
838
+ })).min(1).optional().describe("구버전 영상 전용 입력. 새 제작은 sceneVisuals를 권장합니다. image-motion에서는 사용할 수 없습니다."),
785
839
  featureKey: z.enum(["shortsVideo", "commerceVideo", "musicVideo"]).optional().describe("기본 commerceVideo"),
786
840
  projectTitle: z.string().optional(),
787
841
  aspectRatio: z.string().optional().describe("기본 9:16"),
@@ -808,14 +862,22 @@ export function registerCloudTools(server, config, api) {
808
862
  }).passthrough().optional().describe("measure_channel_spec 이 반환한 fingerprint — 자막 세로 위치·BGM 유무를 참고 채널에 맞춥니다"),
809
863
  transitionMode: z.enum(["fade", "none"]).optional().describe("기본 none(하드컷). fade 는 씬이 맞대기로 붙는 배치에서 컷마다 검은 딥을 만든다 — 명시 요청 시에만 쓸 것"),
810
864
  },
811
- wrapCloudHandler(config, async ({ sceneVideos, featureKey, projectTitle = "aimakeall-mcp", aspectRatio = "9:16", ttsAudioPath = "", bgmAudioPath = "", audioDurationSec = 0, bgmVolume = null, muteVideoAudio = true, titleText = "", subtitleLines, emphasisKeywords, emphasisColor = "", emphasisSizeScale = 0, transitionMode = "none", channelFingerprint = null, productionPath, productionProfile, titleStylePresetId, subtitleStylePresetId, titleStyle, subtitleStyle, titleSegments, overlays, decorationRationale }) => {
865
+ wrapCloudHandler(config, async ({ sceneVideos, sceneVisuals, autoCompose, imageSourceMode, visualMotionMode, featureKey, projectTitle = "aimakeall-mcp", aspectRatio = "9:16", ttsAudioPath = "", bgmAudioPath = "", audioDurationSec = 0, bgmVolume = null, muteVideoAudio = true, titleText = "", subtitleLines, emphasisKeywords, emphasisColor = "", emphasisSizeScale = 0, transitionMode = "none", channelFingerprint = null, productionPath, productionProfile, titleStylePresetId, subtitleStylePresetId, titleStyle, subtitleStyle, titleSegments, overlays, decorationRationale }) => {
812
866
  const context = readProductionContext(config.stateDir, productionPath);
813
867
  productionProfile = context?.productionProfile || productionProfile;
868
+ const visualSettings = resolveVisualSettings(context, { productionProfile, imageSourceMode, visualMotionMode });
869
+ if ((!sceneVisuals?.length && !sceneVideos?.length) || (sceneVisuals?.length && sceneVideos?.length)) {
870
+ return textResult("sceneVisuals 또는 구버전 sceneVideos 중 하나만 전달하세요.", { isError: true });
871
+ }
872
+ if (visualSettings.visualMotionMode === "image-motion" && (sceneVideos?.length || sceneVisuals.some((scene) => scene.kind !== "image"))) {
873
+ return textResult("IMAGE_MOTION_ONLY: 이미지 모션 제작에는 sceneVisuals(kind=image)만 사용할 수 있습니다.", { isError: true });
874
+ }
875
+ autoCompose ??= context?.autoCompose ?? productionProfile?.autoComposeEnabled ?? true;
814
876
  ttsAudioPath ||= context?.tts?.filePath || "";
815
877
  titleText ||= context?.tagline || "";
816
878
  subtitleLines ??= context?.tts?.subtitleLines || [];
817
879
  emphasisKeywords ??= context?.emphasisKeywords || [];
818
- featureKey ||= productionProfile?.categoryId?.endsWith("shorts") ? "shortsVideo" : "commerceVideo";
880
+ featureKey ||= context?.featureKey || (productionProfile?.categoryId?.endsWith("shorts") ? "shortsVideo" : "commerceVideo");
819
881
  channelFingerprint ||= context?.channelFingerprint;
820
882
  // 오디오 확장자 검증 + 합산 크기 예산(서버 413 사전 차단).
821
883
  if (ttsAudioPath) assertAllowedInputFile(ttsAudioPath, new Set([".mp3", ".wav"]), { kind: "TTS 오디오" });
@@ -826,6 +888,7 @@ export function registerCloudTools(server, config, api) {
826
888
  ? audioDurationSec
827
889
  : (ttsAudioPath ? audioDurationSecFromFile(ttsAudioPath) : 0);
828
890
  const body = {
891
+ ...visualSettings, autoCompose,
829
892
  productionProfile, titleStylePresetId, subtitleStylePresetId, titleStyle, subtitleStyle, titleSegments, overlays,
830
893
  aspectRatio,
831
894
  audioDurationSec: resolvedAudioDurationSec,
@@ -834,12 +897,14 @@ export function registerCloudTools(server, config, api) {
834
897
  featureKey,
835
898
  muteVideoAudio,
836
899
  projectTitle,
837
- sceneVideos: sceneVideos.map((scene, index) => ({
900
+ ...(sceneVisuals ? { sceneVisuals: sceneVisuals.map((scene, index) => ({
901
+ ...scene, idx: index, label: scene.label || `scene-${index + 1}`,
902
+ })) } : { sceneVideos: sceneVideos.map((scene, index) => ({
838
903
  durationSec: scene.durationSec,
839
904
  idx: index,
840
905
  label: scene.label || `scene-${index + 1}`,
841
906
  url: scene.url,
842
- })),
907
+ })) }),
843
908
  emphasisColor,
844
909
  emphasisKeywords,
845
910
  channelFingerprint,
@@ -857,7 +922,11 @@ export function registerCloudTools(server, config, api) {
857
922
  if (!response?.ok || !response?.payload) {
858
923
  return textResult(`스티치 실패: ${response?.error || "매니페스트가 비어 있습니다."}`, { isError: true });
859
924
  }
860
- if (context || productionProfile) response.payload.productionContext = { ...context, productionProfile, decorationRationale };
925
+ const autoComposition = response.payload.compositionReport?.autoComposition;
926
+ if (!decorationRationale && autoComposition?.enabled && autoComposition?.omissions?.length) {
927
+ decorationRationale = autoComposition.omissions.map((item) => typeof item === "string" ? item : item?.reason || item?.message || "").filter(Boolean).join(" ").slice(0, 500);
928
+ }
929
+ if (context || productionProfile) response.payload.productionContext = { ...context, productionProfile, visualSettings, decorationRationale };
861
930
  const qualityReview = reviewProductionPayload(response.payload);
862
931
  const handle = savePayloadHandle(config.stateDir, response.payload, "stitch");
863
932
  return jsonResult({
@@ -1157,6 +1226,8 @@ export function registerCloudTools(server, config, api) {
1157
1226
  "plan_music_video",
1158
1227
  "뮤직비디오 씬 플랜을 생성합니다. suno_music_download로 받은 곡의 가사·길이(durationSec)를 입력하세요. 결과 scenes의 imagePrompt는 generate_scene_image로 이어집니다.",
1159
1228
  {
1229
+ ...visualProductionFields,
1230
+ productionProfile: productionFields.productionProfile,
1160
1231
  durationSec: z.number().min(10).describe("곡 길이(초) — suno_music_download 결과의 durationSec"),
1161
1232
  lyrics: z.string().optional().describe("가사 전체 (instrumental이면 생략 가능)"),
1162
1233
  conceptBrief: z.string().optional().describe("영상 컨셉 브리프"),
@@ -1173,12 +1244,14 @@ export function registerCloudTools(server, config, api) {
1173
1244
  })).max(4).optional().describe("등장 캐릭터 — 서버가 씬마다 identity_lock 을 imagePrompt 에 강제 주입해 캐릭터 일관성을 잠급니다"),
1174
1245
  aspectRatio: z.string().optional().describe("기본 9:16"),
1175
1246
  },
1176
- wrapCloudHandler(config, async ({ durationSec, lyrics = "", conceptBrief = "", instrumental = false, sceneSeconds = 5, splitMode = "fixed", videoStylePreset = "seedance-music-video", aspectRatio = "9:16" , characters = [] }) => {
1247
+ wrapCloudHandler(config, async ({ durationSec, lyrics = "", conceptBrief = "", instrumental = false, sceneSeconds = 5, splitMode = "fixed", videoStylePreset = "seedance-music-video", aspectRatio = "9:16" , characters = [], productionProfile, imageSourceMode, visualMotionMode }) => {
1177
1248
  if (!instrumental && !String(lyrics).trim()) {
1178
1249
  return textResult("lyrics를 입력하거나 instrumental=true로 지정하세요.", { isError: true });
1179
1250
  }
1251
+ const visualSettings = resolveVisualSettings(null, { productionProfile, imageSourceMode, visualMotionMode });
1180
1252
  const payload = await api.request("/api/tracker/music-video/plan", {
1181
1253
  body: {
1254
+ ...visualSettings, productionProfile,
1182
1255
  alignedLines: [],
1183
1256
  aspectRatio,
1184
1257
  conceptBrief,
@@ -1197,7 +1270,13 @@ export function registerCloudTools(server, config, api) {
1197
1270
  });
1198
1271
  // usage/원장 이벤트 등 컨텍스트 낭비 필드는 제거하고 기획 본문만 돌려준다.
1199
1272
  const { accountCostEvents, providerUsage, usage, ...plan } = payload && typeof payload === "object" ? payload : {};
1200
- return jsonResult(plan);
1273
+ const scenes = trimPlanScenes(plan.scenes);
1274
+ const productionPath = saveProductionContext(config.stateDir, {
1275
+ topic: conceptBrief, visualSettings, autoCompose: productionProfile?.autoComposeEnabled ?? true,
1276
+ productionProfile: plan.productionProfile || productionProfile,
1277
+ scenes, featureKey: "musicVideo",
1278
+ });
1279
+ return jsonResult({ ...plan, scenes, ...visualSettings, productionPath });
1201
1280
  }),
1202
1281
  );
1203
1282
 
@@ -1314,17 +1393,31 @@ export function registerCloudTools(server, config, api) {
1314
1393
  endMs: z.number().describe("원본에서 자를 끝(ms)"),
1315
1394
  })).min(1).max(6).describe("이 행에 이어 붙일 원본 구간들"),
1316
1395
  })).min(1).max(40).describe("타임라인 행 — 순서대로 이어 붙습니다"),
1317
- provider: z.enum(["typecast", "elevenlabs", "supertone"]).optional().describe("TTS 공급자, 기본 typecast (recommend_voice 로 보이스를 먼저 고르세요)"),
1396
+ provider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("자동: Typecast→ElevenLabs→Supertone→Edge→Supertonic2. recommend_voice로 보이스를 먼저 고르세요."),
1318
1397
  voiceId: z.string().optional().describe("보이스 ID (생략 시 공급자 기본)"),
1319
1398
  aspectRatio: z.enum(["9:16", "16:9", "1:1"]).optional().describe("기본 9:16"),
1320
1399
  subtitleCropMode: z.enum(["off", "always", "auto"]).optional().describe("원본 박힌 자막 하단 크롭 — auto 는 burnedSubtitle.present 소스만"),
1321
1400
  },
1322
- wrapCloudHandler(config, async ({ title = "", sourceVideos, rows, provider = "typecast", voiceId = "", aspectRatio = "9:16", subtitleCropMode = "off" }) => {
1401
+ wrapCloudHandler(config, async ({ title = "", sourceVideos, rows, provider, voiceId = "", aspectRatio = "9:16", subtitleCropMode = "off" }) => {
1323
1402
  const sourceIds = new Set(sourceVideos.map((source) => source.id));
1324
1403
  const badRef = rows.flatMap((row) => row.clipRefs).find((ref) => !sourceIds.has(ref.videoId));
1325
1404
  if (badRef) {
1326
1405
  return textResult(`clipRefs 의 videoId "${badRef.videoId}" 가 sourceVideos 에 없습니다.`, { isError: true });
1327
1406
  }
1407
+ const hasNarration = rows.some((row) => ["N", "SN"].includes(row.mode) && row.audioContent?.trim());
1408
+ const selection = hasNarration ? await resolveTtsSelection(api, config, { provider }) : { provider: "typecast", executionTarget: "server" };
1409
+ provider = selection.provider;
1410
+ let preSynthesizedNarrations;
1411
+ if (selection.executionTarget === "companion") {
1412
+ const narrations = {};
1413
+ for (const [rowIndex, row] of rows.entries()) {
1414
+ if (!["N", "SN"].includes(row.mode) || !row.audioContent?.trim()) continue;
1415
+ const text = row.audioContent.trim();
1416
+ const result = await requestTtsBinary(api, config, selection, buildTtsRequestBody(provider, { text, voice: voiceId }));
1417
+ narrations[rowIndex] = { dataUrl: `data:${result.contentType};base64,${result.buffer.toString("base64")}`, charCount: Array.from(text).length, voiceId };
1418
+ }
1419
+ preSynthesizedNarrations = { narrations, attempted: Object.keys(narrations).length, succeeded: Object.keys(narrations).length, lastError: "" };
1420
+ }
1328
1421
  const version = {
1329
1422
  id: "mcp-remake",
1330
1423
  title,
@@ -1342,6 +1435,7 @@ export function registerCloudTools(server, config, api) {
1342
1435
  aspectRatio,
1343
1436
  costEventId: createUsageEventId("mcp-remake-tts"),
1344
1437
  provider,
1438
+ ...(preSynthesizedNarrations ? { preSynthesizedNarrations } : {}),
1345
1439
  subtitleCropMode,
1346
1440
  version,
1347
1441
  voiceId,
@@ -1370,27 +1464,47 @@ export function registerCloudTools(server, config, api) {
1370
1464
  // ── 시간 동기 자막 TTS — 쇼츠 하단자막용 ─────────────────────────────────────
1371
1465
  server.tool(
1372
1466
  "tts_narration_with_captions",
1373
- "ElevenLabs TTS를 char-level 타이밍과 함께 합성해 mp3 파일 + 시간 동기 자막 라인(subtitleLines)을 만듭니다. 쇼츠 하단자막이 필요하면 tts_narration 대신 이걸 쓰고, 결과 subtitleLines·durationSec·filePath를 stitch_timeline에 그대로 넘기세요 (웹 쇼츠 스튜디오와 동일한 정렬 방식). 사용자가 보이스를 지정하지 않았다면 먼저 recommend_voice(provider=elevenlabs)로 자동 선정하세요.",
1467
+ "TTS 오디오와 자막을 만듭니다. 자동 순서: Typecast→ElevenLabs→Supertone→Edge→Supertonic2, 명시 보이스/공급자는 유지. ElevenLabs는 실제 char-level 타이밍, 나머지는 실측 오디오 길이에 비례한 추정 자막(timingSource 확인)입니다. 먼저 recommend_voice로 제작 톤에 맞는 보이스를 고르세요. 자동 추가 STT 과금이나 다른 공급자로의 재합성은 하지 않습니다.",
1374
1468
  {
1375
1469
  ...productionFields,
1376
1470
  text: z.string().optional().describe("생략 시 제작 핸들의 기획 대본 그대로 합성. 사실·인용을 바꾸거나 뉴스체로 재작성하지 마세요."),
1471
+ provider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("기본 자동 선택 또는 제작 핸들에 선정된 공급자"),
1377
1472
  voice: z.string().optional().describe("보이스 이름 또는 ID (생략 시 기본 보이스)"),
1378
1473
  speed: z.number().optional().describe("기본 1 (0.7~1.2)"),
1379
1474
  modelId: z.string().optional(), stability: z.number().min(0).max(1).optional(),
1380
1475
  style: z.number().min(0).max(1).optional(), similarityBoost: z.number().min(0).max(1).optional(), speakerBoost: z.boolean().optional(),
1381
1476
  },
1382
- wrapCloudHandler(config, async ({ text, voice = "", speed, productionPath, productionProfile, modelId, stability, style, similarityBoost, speakerBoost }) => {
1477
+ wrapCloudHandler(config, async ({ text, provider, voice = "", speed, productionPath, productionProfile, modelId, stability, style, similarityBoost, speakerBoost }) => {
1383
1478
  const context = readProductionContext(config.stateDir, productionPath);
1384
1479
  productionProfile = context?.productionProfile || productionProfile;
1385
1480
  text ??= context?.narrationScript;
1386
- if (!voice && context?.voice?.provider && context.voice.provider !== "elevenlabs") {
1387
- return textResult("이 제작 핸들의 보이스는 ElevenLabs용이 아닙니다. recommend_voice(provider=elevenlabs)로 다시 선정하거나 ElevenLabs voice를 명시하세요.", { isError: true });
1481
+ const selection = await resolveTtsSelection(api, config, { provider, context });
1482
+ provider = selection.provider;
1483
+ if (!voice && context?.voice?.provider && context.voice.provider !== provider) {
1484
+ return textResult("보이스 공급자가 다릅니다. 선택한 provider용 voice를 다시 선정하세요.", { isError: true });
1388
1485
  }
1389
1486
  voice ||= context?.voice?.voiceId || "";
1390
1487
  if (!text?.trim()) return textResult("text 또는 기획 대본이 있는 productionPath가 필요합니다.", { isError: true });
1391
1488
  if (productionProfile?.deliveryStyle === "casual-social" && !voice) {
1392
1489
  return textResult("커뮤니티형 음성을 위해 먼저 recommend_voice에 productionPath를 전달하고, 반환된 productionPath 또는 보이스 ID를 사용하세요.", { isError: true });
1393
1490
  }
1491
+ if (provider !== "elevenlabs") {
1492
+ const body = buildTtsRequestBody(provider, { productionProfile, text, voice, speed, modelId, costEventId: createUsageEventId(`mcp-tts-${provider}-captions`) });
1493
+ const result = await requestTtsBinary(api, config, selection, body);
1494
+ const saved = saveMediaBuffer(config.stateDir, provider === "supertonic2" ? "narration-captions.wav" : "narration-captions.mp3", result.buffer, { contentType: result.contentType });
1495
+ const durationSec = audioDurationSecFromFile(saved.filePath);
1496
+ const subtitleLines = estimateTtsSubtitleLines(text, durationSec);
1497
+ const timingSource = "estimated-from-audio-duration";
1498
+ const updatedProductionPath = saveProductionContext(config.stateDir, { ...context, productionProfile, narrationScript: text,
1499
+ tts: { provider, durationSec, filePath: saved.filePath, subtitleLines, voiceId: voice, timingSource },
1500
+ });
1501
+ return jsonResult({ productionPath: updatedProductionPath, productionProfile, provider, executionTarget: selection.executionTarget,
1502
+ durationSec, durationWarning: durationReview(productionProfile?.targetDurationSec, durationSec), filePath: saved.filePath,
1503
+ lineCount: subtitleLines.length, subtitleLines, timingSource,
1504
+ warnings: [durationSec > 0 ? "자막은 실측 오디오 길이 기준 추정입니다. 최종 렌더 전 발화와 동기를 확인하세요." : "오디오 길이를 측정하지 못해 자막을 만들지 않았습니다. 저장된 오디오를 확인하세요. 자동 재합성하지 않습니다."],
1505
+ next: "이 productionPath를 stitch_timeline에 전달하세요. 추정 자막의 발화 동기는 최종 미리보기에서 확인하세요.",
1506
+ });
1507
+ }
1394
1508
  const payload = await api.request("/api/tracker/tts/elevenlabs-with-timestamps", {
1395
1509
  body: {
1396
1510
  productionProfile, modelId, stability, style, similarityBoost, speakerBoost,
@@ -1421,10 +1535,12 @@ export function registerCloudTools(server, config, api) {
1421
1535
  : audioDurationSecFromFile(saved.filePath);
1422
1536
  const updatedProductionPath = saveProductionContext(config.stateDir, {
1423
1537
  ...context, productionProfile, narrationScript: text,
1424
- tts: { durationSec, filePath: saved.filePath, subtitleLines, voiceId: payload?.voiceId || voice, deliverySettings: payload?.deliverySettings },
1538
+ tts: { provider, durationSec, filePath: saved.filePath, subtitleLines, voiceId: payload?.voiceId || voice, deliverySettings: payload?.deliverySettings, timingSource: "provider-character-alignment" },
1425
1539
  });
1426
1540
  return jsonResult({
1427
1541
  productionPath: updatedProductionPath,
1542
+ provider,
1543
+ timingSource: "provider-character-alignment",
1428
1544
  productionProfile,
1429
1545
  deliverySettings: payload?.deliverySettings,
1430
1546
  durationWarning: durationReview(productionProfile?.targetDurationSec, durationSec),
@@ -1442,10 +1558,12 @@ export function registerCloudTools(server, config, api) {
1442
1558
  "tts_list_voices",
1443
1559
  "TTS 보이스 목록을 조회합니다. 사용자가 목록에서 직접 고르고 싶어할 때 쓰세요 — 주제·샘플 영상 기반 자동 선정은 recommend_voice가 담당합니다. (voice를 생략하면 서버가 기본 보이스로 합성합니다)",
1444
1560
  {
1445
- provider: z.enum(["typecast", "elevenlabs"]).optional().describe("기본 typecast"),
1561
+ provider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("기본 자동: Typecast→ElevenLabs→Supertone→Edge→Supertonic2"),
1446
1562
  },
1447
- wrapCloudHandler(config, async ({ provider = "typecast" }) => {
1448
- const payload = await api.request(`/api/tracker/tts/${provider}/voices`, { method: "GET", timeoutMs: 30_000 });
1563
+ wrapCloudHandler(config, async ({ provider }) => {
1564
+ const selection = await resolveTtsSelection(api, config, { provider });
1565
+ provider = selection.provider;
1566
+ const payload = selection.localVoices ? { voices: selection.localVoices } : await api.request(`/api/tracker/tts/${provider}/voices`, { method: "GET", timeoutMs: 30_000 });
1449
1567
  // ElevenLabs 는 gender/age 가 labels 객체 안에 있다 — 두 공급자 형태 모두 지원.
1450
1568
  const voices = (Array.isArray(payload?.voices) ? payload.voices : []).slice(0, 80).map((voice) => ({
1451
1569
  age: voice?.age || voice?.labels?.age,
@@ -1455,7 +1573,7 @@ export function registerCloudTools(server, config, api) {
1455
1573
  nameKo: voice?.voiceNameKo,
1456
1574
  voiceId: voice?.voiceId,
1457
1575
  }));
1458
- return jsonResult({ count: voices.length, provider, voices });
1576
+ return jsonResult({ count: voices.length, provider, executionTarget: selection.executionTarget, voices });
1459
1577
  }),
1460
1578
  );
1461
1579
 
@@ -1467,13 +1585,14 @@ export function registerCloudTools(server, config, api) {
1467
1585
  topic: z.string().optional().describe("영상 주제·타깃 시청층 (예: 시니어 건강 정보, 커머스 쇼츠)"),
1468
1586
  scriptExcerpt: z.string().optional().describe("대본 앞부분 발췌 (선택)"),
1469
1587
  sampleVideoUrl: z.string().optional().describe("사용자가 샘플로 지정한 유튜브 URL 또는 11자 영상 ID"),
1470
- provider: z.enum(["typecast", "elevenlabs"]).optional().describe("기본 typecast"),
1588
+ provider: z.enum(["auto", ...TTS_PROVIDERS]).optional().describe("기본 자동: Typecast→ElevenLabs→Supertone→Edge→Supertonic2. 키가 없으면 다음 후보, 합성 실패 후 자동 재과금 없음"),
1471
1589
  sampleVoiceHints: z.string().optional().describe("사용자가 말한 보이스 요구 (예: 차분한 중년 남성)"),
1472
1590
  },
1473
1591
  wrapCloudHandler(config, async ({ topic = "", scriptExcerpt = "", sampleVideoUrl = "", provider, sampleVoiceHints = "", productionPath, productionProfile }) => {
1474
1592
  const context = readProductionContext(config.stateDir, productionPath);
1475
1593
  productionProfile = context?.productionProfile || productionProfile;
1476
- provider ||= productionProfile ? "elevenlabs" : "typecast";
1594
+ const selection = await resolveTtsSelection(api, config, { provider, context });
1595
+ provider = selection.provider;
1477
1596
  topic ||= context?.topic || "";
1478
1597
  scriptExcerpt ||= context?.narrationScript?.slice(0, 2000) || "";
1479
1598
  const notes = [];
@@ -1504,7 +1623,7 @@ export function registerCloudTools(server, config, api) {
1504
1623
  return textResult("topic, scriptExcerpt, sampleVoiceHints, sampleVideoUrl 중 하나는 필요합니다.", { isError: true });
1505
1624
  }
1506
1625
  const payload = await api.request("/api/tracker/tts/recommend-voice", {
1507
- body: { provider, sampleAudioBase64, sampleVoiceHints, scriptExcerpt, topic, productionProfile },
1626
+ body: { provider, sampleAudioBase64, sampleVoiceHints, scriptExcerpt, topic, productionProfile, ...(selection.localVoices ? { localVoiceCatalog: selection.localVoices } : {}) },
1508
1627
  method: "POST",
1509
1628
  timeoutMs: 180_000,
1510
1629
  });
@@ -66,7 +66,7 @@ export async function recommendComposerOverlays(api, { subtitleLines, context =
66
66
 
67
67
  export function registerComposerTools(server, config, api) {
68
68
  server.tool("list_composer_presets",
69
- "Cut Composer의 실제 제목·자막 디자인 프리셋과 애니메이션을 조회합니다(유료 생성 없음). 반환된 preset ID/style을 stitch_timeline의 제목·자막 설정에 사용하세요. 단일 기본 폰트 대신 요청 톤에 맞는 서로 다른 제목/자막 프리셋을 선택하고 가독성을 검토하세요.",
69
+ "Cut Composer의 실제 제목·자막 디자인, 등장/강조 애니메이션과 이미지 인스펙터 imageMotions를 조회합니다(유료 생성 없음). imageMotions의 id를 stitch_timeline.sceneVisuals[].imageMotionPreset에, 강도(0~100)를 imageMotionIntensity에 사용하세요. 제목/본문 계층을 구분하고 가독성을 검토하세요.",
70
70
  { query: z.string().max(500).optional(), category: z.string().max(80).optional(), limit: z.number().int().min(1).max(100).optional() },
71
71
  READ_ONLY, withCatalogAuth(config, (args) => api.request(PRESETS_PATH, { method: "POST", body: args, timeoutMs: 30_000 })));
72
72
 
package/lib/config.mjs CHANGED
@@ -2,7 +2,7 @@ import { homedir } from "node:os";
2
2
  import path from "node:path";
3
3
 
4
4
  // 프록시 버전 — 서버가 X-AImakeAll-MCP-Version 으로 하한을 강제(426)할 수 있다.
5
- export const MCP_PROXY_VERSION = "0.11.0";
5
+ export const MCP_PROXY_VERSION = "0.12.0";
6
6
 
7
7
  export const DEFAULT_API_BASE = "https://aimakeall.com";
8
8
  export const DEFAULT_COMPANION_URL = "http://127.0.0.1:9876";
@@ -43,6 +43,7 @@ export function reviewProductionPayload(payload, { measured = null } = {}) {
43
43
  const emphasisCount = textClips.filter(c => c.textSegments?.some(s => s.color || s.sizeScale > 1)).length;
44
44
  const animationCount = clips.filter(c => [c.animationIntroPreset,c.animationEmphasisPreset,c.animationOutroPreset].some(v => v && v !== "none")).length;
45
45
  const report = payload?.compositionReport || {};
46
+ const imageMotionCount = clips.filter(c => c.imageMotionPreset && Number(c.imageMotionIntensity ?? 50) > 0).length;
46
47
  if (profile?.deliveryStyle === "casual-social") {
47
48
  if (!emphasisCount) warnings.push({ code: "NO_EMPHASIS", message: "강조 단어가 실제 자막과 일치하는지 확인하세요." });
48
49
  if (!animationCount) warnings.push({ code: "NO_MOTION", message: "자막·타이틀의 등장/강조 애니메이션이 없습니다." });
@@ -54,7 +55,7 @@ export function reviewProductionPayload(payload, { measured = null } = {}) {
54
55
  return {
55
56
  pass: errors.length === 0, errors, warnings,
56
57
  targetDurationSec: profile?.targetDurationSec || null, actualDurationSec: actual,
57
- emphasisCount, animationCount, overlayCount: report.overlayCount || 0,
58
+ emphasisCount, animationCount, imageMotionCount, overlayCount: report.overlayCount || 0,
58
59
  visualReviewRequired: true, listeningReviewRequired: true,
59
60
  note: "구조 검증만 수행했습니다. 최종 영상의 글자 크기·색상·가독성과 음성 말투를 직접 확인하기 전에는 품질 검증 완료로 보고하지 마세요.",
60
61
  };
@@ -1,13 +1,24 @@
1
1
  import { z } from "zod";
2
2
 
3
+ export const visualProductionFields = {
4
+ imageSourceMode: z.enum(["web", "ai", "mixed"]).optional().describe("장면 이미지 준비: web=스톡 검색만(없으면 중단), ai=AI 생성만, mixed=웹 우선 후 없을 때 AI 생성. 기본 ai. 편집용 아이콘/이모지는 별도입니다."),
5
+ visualMotionMode: z.enum(["image-motion", "i2v"]).optional().describe("image-motion=이미지 인스펙터 모션으로 편집(I2V 호출 금지), i2v=AI 영상 생성. 기본 i2v."),
6
+ };
7
+
3
8
  export const productionFields = {
4
9
  productionPath: z.string().optional().describe("plan_shorts_video / TTS가 반환한 제작 핸들. 이후 단계에도 반드시 이어서 전달해 스타일·대본·목표 길이를 유지하세요."),
5
10
  productionProfile: z.object({
6
11
  version: z.literal(1), categoryId: z.string(), tone: z.string().optional(),
7
12
  deliveryStyle: z.enum(["casual-social", "standard", "news-briefing"]),
8
- targetDurationSec: z.number().positive().max(180),
13
+ targetDurationSec: z.number().positive().max(3600),
9
14
  titleStylePresetId: z.string().optional(), subtitleStylePresetId: z.string().optional(),
10
- }).optional().describe("제작 스타일 명시. 바이럴/커뮤니티는 casual-social. 핸들이 있으면 저장된 프로필을 사용합니다."),
15
+ ...visualProductionFields,
16
+ autoComposeEnabled: z.boolean().optional().describe("장면에 맞는 이미지 모션·자막 계층·라이브러리 소재 자동 편집"),
17
+ }).superRefine((profile, ctx) => {
18
+ if (["community-shorts", "viral-shorts"].includes(profile.categoryId) && profile.targetDurationSec > 180) {
19
+ ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["targetDurationSec"], message: "커뮤니티/바이럴 쇼츠는 최대 180초입니다." });
20
+ }
21
+ }).optional().describe("제작 스타일 명시. 바이럴/커뮤니티는 casual-social, 최대 180초; 일반/뮤직비디오는 최대 3600초. 핸들이 있으면 저장된 프로필을 사용합니다."),
11
22
  };
12
23
  export const composerTextStyle = z.record(z.any()).describe("list_composer_presets의 실제 style 속성. fontFamily, 색상, 배경, 외곽선, animationIntroPreset/EmphasisPreset/OutroPreset 등을 지원하며 서버가 검증합니다.");
13
24
  export const composerOverlay = z.object({
@@ -19,4 +30,14 @@ export const composerOverlay = z.object({
19
30
  widthPct: z.number().positive().max(100).optional(), heightPct: z.number().positive().max(100).optional(),
20
31
  opacity: z.number().optional(), volume: z.number().optional(),
21
32
  animationIntroPreset: z.string().optional(), animationEmphasisPreset: z.string().optional(), animationOutroPreset: z.string().optional(),
33
+ imageMotionPreset: z.string().optional(), imageMotionIntensity: z.number().min(0).max(100).optional(),
34
+ });
35
+
36
+ export const sceneVisual = z.object({
37
+ kind: z.enum(["image", "video"]), idx: z.number().int().min(0), url: z.string().min(1),
38
+ durationSec: z.number().positive(), label: z.string().optional(), role: z.string().optional(),
39
+ fitMode: z.enum(["cover", "contain"]).optional().describe("웹 이미지 기본 contain으로 원본 내용을 보존. cover는 화면 채우기 크롭."),
40
+ imageMotionPreset: z.string().optional(), imageMotionIntensity: z.number().min(0).max(100).optional(),
41
+ source: z.enum(["web", "ai"]).optional(), sourcePage: z.string().optional(), license: z.string().optional(),
42
+ attribution: z.union([z.string(), z.object({ provider: z.string().optional(), author: z.string().optional(), url: z.string().optional() })]).optional(),
22
43
  });
@@ -0,0 +1,74 @@
1
+ export const TTS_PROVIDERS = Object.freeze(["typecast", "elevenlabs", "supertone", "edge", "supertonic2"]);
2
+
3
+ export async function resolveTtsSelection(api, config, { provider = "auto", context = null } = {}) {
4
+ provider = provider === "auto" || !provider ? context?.voice?.provider || "auto" : provider;
5
+ if (provider === "edge-tts") provider = "edge";
6
+ if (provider !== "auto" && !TTS_PROVIDERS.includes(provider)) throw new Error("지원하지 않는 TTS provider입니다.");
7
+ // Explicit paid providers and saved castings must not silently change voice.
8
+ if (provider !== "auto" && provider !== "edge") return { provider, executionTarget: "server" };
9
+ const availability = await api.request("/api/tracker/tts/providers", { method: "GET", timeoutMs: 15_000 });
10
+ if (!Array.isArray(availability?.providers)) throw new Error("TTS 공급자 설정 응답이 없습니다. 자동 선택을 중단합니다.");
11
+ const providers = availability.providers;
12
+ const available = (id) => providers.some((entry) => entry.id === id && entry.available === true);
13
+ if (provider === "auto") {
14
+ const paid = TTS_PROVIDERS.slice(0, 3).find(available);
15
+ if (paid) return { provider: paid, executionTarget: "server" };
16
+ }
17
+ if (available("edge")) return { provider: "edge", executionTarget: "server" };
18
+ if (config.companionUrl) {
19
+ try {
20
+ const response = await fetch(`${config.companionUrl}/api/edge-tts/voices`, { signal: AbortSignal.timeout(15_000) });
21
+ const payload = await response.json();
22
+ if (response.ok && Array.isArray(payload?.voices) && payload.voices.length) return { provider: "edge", executionTarget: "companion", localVoices: payload.voices };
23
+ } catch { /* A read-only free-runtime probe can fail without triggering billing. */ }
24
+ }
25
+ if (provider === "auto" && available("supertonic2")) return { provider: "supertonic2", executionTarget: "server" };
26
+ throw new Error(provider === "edge" ? "선택한 Edge TTS 런타임을 사용할 수 없습니다. 로컬 처리기를 확인하세요." : "사용 가능한 TTS가 없습니다. API 키 또는 Edge/Supertonic2 런타임을 확인하세요.");
27
+ }
28
+
29
+ export function buildTtsRequestBody(provider, args = {}) {
30
+ const { productionProfile, text, voice = "", speed, costEventId, modelId, stability, style, similarityBoost, speakerBoost } = args;
31
+ const resolvedSpeed = speed ?? (productionProfile?.deliveryStyle === "casual-social" ? 1.1 : undefined);
32
+ const common = { productionProfile, text, voice, voiceId: voice, speed: resolvedSpeed, costEventId, fileName: "mcp-narration", modelId };
33
+ if (provider === "typecast") return { ...common, language: "kor", emotion: "neutral", smartEmotion: true };
34
+ if (provider === "elevenlabs") return { ...common, languageCode: "ko", stability, style, similarityBoost, speakerBoost };
35
+ if (provider === "edge") {
36
+ const rate = Math.round(((resolvedSpeed ?? 1) - 1) * 100);
37
+ return { ...common, rate: `${rate >= 0 ? "+" : ""}${rate}%`, pitch: "+0Hz" };
38
+ }
39
+ // Numeric ElevenLabs style is not a valid Supertone style ID.
40
+ return { ...common, language: "ko", lang: "ko" };
41
+ }
42
+
43
+ export async function requestTtsBinary(api, config, selection, body) {
44
+ if (selection.executionTarget !== "companion") return api.requestBinary(`/api/tracker/tts/${selection.provider}`, { body, method: "POST", timeoutMs: selection.provider === "supertonic2" ? 600_000 : 180_000 });
45
+ const voiceId = body.voiceId || selection.localVoices?.find((voice) => voice.voiceId === "ko-KR-SunHiNeural")?.voiceId || selection.localVoices?.find((voice) => String(voice.locale || "").startsWith("ko"))?.voiceId;
46
+ if (!voiceId) throw new Error("사용 가능한 한국어 Edge 보이스를 찾지 못했습니다. 보이스를 직접 지정하세요.");
47
+ const response = await fetch(`${config.companionUrl}/api/edge-tts/synthesize`, { body: JSON.stringify({ ...body, voiceId }), headers: { "Content-Type": "application/json" }, method: "POST", signal: AbortSignal.timeout(180_000) });
48
+ if (!response.ok) throw new Error(`Edge TTS 합성 실패 (HTTP ${response.status}). 다른 공급자로 자동 재합성하지 않습니다.`);
49
+ return { buffer: Buffer.from(await response.arrayBuffer()), contentType: response.headers.get("content-type") || "audio/mpeg", headers: response.headers };
50
+ }
51
+
52
+ // Other providers currently do not expose char-level timestamps. These are
53
+ // explicitly approximate captions, without another paid STT/synthesis call.
54
+ export function estimateTtsSubtitleLines(text, durationSec, maxCharacters = 30) {
55
+ if (!Number.isFinite(durationSec) || durationSec <= 0) return [];
56
+ const lines = [];
57
+ for (const sentence of String(text || "").split(/(?<=[.!?。!?])\s+|\n+/u).map((part) => part.trim()).filter(Boolean)) {
58
+ let current = "";
59
+ for (const word of sentence.split(/\s+/u)) {
60
+ if (current && Array.from(`${current} ${word}`).length > maxCharacters) { lines.push(current); current = ""; }
61
+ const chars = Array.from(word);
62
+ while (chars.length > maxCharacters) { if (current) { lines.push(current); current = ""; } lines.push(chars.splice(0, maxCharacters).join("")); }
63
+ current = [current, chars.join("")].filter(Boolean).join(" ");
64
+ }
65
+ if (current) lines.push(current);
66
+ }
67
+ const total = lines.reduce((sum, line) => sum + Array.from(line).length, 0);
68
+ let cursor = 0;
69
+ return lines.map((line, index) => {
70
+ const startSec = cursor;
71
+ cursor += durationSec * Array.from(line).length / total;
72
+ return { text: line, startSec, endSec: index === lines.length - 1 ? durationSec : cursor };
73
+ });
74
+ }
@@ -0,0 +1,56 @@
1
+ // Kept dependency-free: the published MCP package must not import the repo's UI.
2
+ export const STOCK_IMAGE_SEARCH_PATH = "/api/tracker/media-search/stock";
3
+
4
+ export function resolveVisualSettings(context, input = {}) {
5
+ const stored = { ...context?.productionProfile, ...context?.visualSettings };
6
+ const result = {
7
+ imageSourceMode: stored.imageSourceMode ?? input.productionProfile?.imageSourceMode ?? input.imageSourceMode ?? "ai",
8
+ visualMotionMode: stored.visualMotionMode ?? input.productionProfile?.visualMotionMode ?? input.visualMotionMode ?? "i2v",
9
+ };
10
+ if (!["web", "ai", "mixed"].includes(result.imageSourceMode)) throw new Error("지원하지 않는 장면 이미지 준비 방식입니다.");
11
+ if (!["image-motion", "i2v"].includes(result.visualMotionMode)) throw new Error("지원하지 않는 영상 표현 방식입니다.");
12
+ return result;
13
+ }
14
+
15
+ export function assertI2vAllowed(settings) {
16
+ if (settings.visualMotionMode === "image-motion") {
17
+ throw new Error("IMAGE_MOTION_ONLY: 이미지 모션 편집을 선택한 제작입니다. 유료 I2V 프롬프트/영상 생성은 실행하지 않습니다. 이미지 URL을 stitch_timeline.sceneVisuals(kind=image)에 전달하세요.");
18
+ }
19
+ }
20
+
21
+ export function sceneWebSearchQuery(scene = {}) {
22
+ return String(scene.webSearchQuery || scene.imageSearchQuery || scene.label || scene.imagePrompt || "")
23
+ .replace(/\s+/g, " ").trim().slice(0, 160);
24
+ }
25
+
26
+ // Stock provider response fields are untrusted. Restrict automatic selection to
27
+ // the providers' known HTTPS image CDNs, not arbitrary URLs, redirects or IPs.
28
+ export function isEligibleStockImage(item) {
29
+ if (!item || item.contentKind === "video" || /^(mp4|webm|mov|m4v)$/i.test(String(item.format || ""))) return false;
30
+ try {
31
+ const url = new URL(String(item.url || ""));
32
+ if (url.protocol !== "https:" || url.username || url.password || (url.port && url.port !== "443")) return false;
33
+ if (item.source === "pexels") return url.hostname === "images.pexels.com";
34
+ if (item.source === "pixabay") return url.hostname === "cdn.pixabay.com" || url.hostname === "pixabay.com";
35
+ return false;
36
+ } catch { return false; }
37
+ }
38
+
39
+ export async function findSceneWebImage(api, { query, excludeImageUrls = [] } = {}) {
40
+ const cleanQuery = sceneWebSearchQuery({ webSearchQuery: query });
41
+ if (!cleanQuery) throw new Error("웹 이미지 검색에는 장면의 짧은 검색어(webSearchQuery)가 필요합니다.");
42
+ const payload = await api.request(STOCK_IMAGE_SEARCH_PATH, {
43
+ method: "POST", body: { query: cleanQuery, sources: ["pexels", "pixabay"], page: 1, perPage: 24 }, timeoutMs: 45_000,
44
+ });
45
+ if (payload?.ok === false) throw new Error("웹 이미지 검색에 실패했습니다. 검색 키와 공급자 연결을 확인하세요. AI 생성으로 자동 전환하지 않았습니다.");
46
+ const excluded = new Set(excludeImageUrls);
47
+ const item = (Array.isArray(payload?.items) ? payload.items : []).find((entry) => isEligibleStockImage(entry) && !excluded.has(entry.url));
48
+ if (!item) return null;
49
+ return {
50
+ kind: "image", url: item.url, source: "web", stockProvider: item.source,
51
+ sourcePage: item.sourcePage || item.attribution?.url || "",
52
+ attribution: item.attribution || item.sourceLabel || item.source,
53
+ license: item.license || "", label: item.title || cleanQuery,
54
+ webSearchQuery: cleanQuery,
55
+ };
56
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "aimakeall-mcp",
3
- "version": "0.11.0",
3
+ "version": "0.12.0",
4
4
  "description": "AImakeAll MCP 서버 — Claude Code/Codex에서 자연어로 영상 기획·생성·렌더·퍼블리시 (렌더는 로컬 컴패니언)",
5
5
  "type": "module",
6
6
  "bin": {