free-short-video 6.3.0 → 6.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/.env.example CHANGED
@@ -77,3 +77,13 @@ PORT=8765
77
77
 
78
78
  # 提示词语言(zh/en,影响 LLM meta-prompt 语言)
79
79
  # PROMPT_LANGUAGE=zh
80
+
81
+
82
+ # ── 7. CORS 跨源白名单(可选,供独立本地伴侣工具调用本服务 API)──────────
83
+ # 场景:本地跑一个独立的简化前端(如 agnes-simple-ui,监听 :8787),
84
+ # 从浏览器直接调用本服务的 /api/* 接口。设置后服务仅允许列出的源跨源调用。
85
+ # 默认空 = 不启用 CORS(攻击面不变);同源页面不受影响。
86
+ # AGNES_CORS_ORIGINS=http://localhost:8787,http://127.0.0.1:3000
87
+ #
88
+ # 显式禁用(即使设置了 origins 也不启用中间件):
89
+ # AGNES_CORS_ENABLED=false
package/README.md CHANGED
@@ -1,36 +1,34 @@
1
1
  ---
2
2
 
3
- # What's New in v6.3.0
3
+ # What's New in v6.4.0
4
4
 
5
5
  ## What's New
6
6
 
7
7
  ### Features & Improvements
8
8
 
9
- - **Complete v6 optimization roadmap (29/29 items)** — every batch of the v6 roadmap is now shipped:
10
- - **Performance (batch 2)**: the final compositing chain is now ffmpeg-based — identical-parameter scene concatenation uses `-c copy`, audio alignment/volume/silence-padding merge into a single filter pass, and subtitles render through the ASS path with per-entry styles (`AGNES_SUBTITLE_ASS`, with automatic fallback to the moviepy path). Poetry videos compose all scenes in one pass instead of re-encoding per scene. A dedicated encoding thread pool isolates heavy ffmpeg/moviepy work from API requests, and the token-bucket rate limiter gained a native async path so stopping a task during rate-limit waits is instant.
11
- - **Reliability & engineering (batch 1)**: task state follows a single-writer principle with per-task locking, resume supports persisted word-level TTS cues (no re-synthesis on resume), video polling is adaptive and multi-scene waits run concurrently, task listing is indexed with `limit/offset/status` pagination, stale artifacts/error logs are governed, and the frontend stops polling in background tabs with exponential backoff and a connection-loss banner.
12
- - **Frontend & i18n**: translations are split into per-language lazy-loaded chunks — the first-screen JS bundle drops from ~721 kB to ~305 kB (gzip 226 kB → 97 kB, **-58%**). Form submission/confirm/toast flows were unified into shared composables, mobile layout, focus-trap modals, `prefers-reduced-motion` and form drafts were added.
13
- - **Observability & ops (batch 3)**: new `GET /api/health` and `GET /api/metrics` endpoints, optional rotating file logging (`AGNES_LOG_FILE`), and a Docker `HEALTHCHECK`. Runtime settings are now converged through typed `pydantic-settings` (with `.env` support) so concurrency limits scale dynamically with API-key count.
14
- - **Immediate defect fixes (batch 0)**: stop now cancels instantly without retry backoff, event-loop blocking (watermark re-encode, sync downloads) is moved off the loop, multi-key delete works correctly, a frontend `v-html` XSS vector is closed, and image generation got a duplicate-submit guard.
15
- - **Full 22-language support incl. Arabic** — the UI already had 22 languages; this release completes the voice catalog for all of them. Arabic UI is fully supported (PR #32), and 8 UI languages (Turkish, Vietnamese, Thai, Tagalog, Hindi, Persian, Bengali, Urdu) now have edge_tts voice groupings with native-voice name display, script-detection regexes (Thai/Devanagari/Bengali) and per-script subtitle font fallback (new bundled Noto fonts; Persian/Urdu reuse the Arabic reshape+bidi pipeline).
16
- - **Transparent analytics disclosure & privacy controls** — the settings panel now shows a clear, collapsible privacy card listing exactly what usage statistics are reported (and what is never uploaded: prompts, manuscripts, poems, API keys and reference images are redacted before reporting). Analytics can be turned off entirely from the panel.
17
- - **Complete error tracebacks in the feedback report** — pipeline failures now persist the full `traceback` into the task state; the diagnostics endpoint and the in-app feedback report include it, so you can paste complete error details (e.g. environment-level `[WinError 2]`) into GitHub issues without checking the server console.
9
+ - **Preview ("dry-run") endpoints for creative and manuscript tasks** — `POST /api/creative/preview-script` and `POST /api/manuscript/preview-split` let you inspect the generated script or manuscript segments before submitting a full task. Synchronous, no task created. Preview calls share the Chat API rate limiter and have a lightweight in-process concurrency cap (429 + Retry-After) to prevent abuse. (Credit: @Khaled97Sho, PR #33)
10
+ - **Arabic tashkeel for TTS narration** — creative and manuscript pipelines now automatically apply Arabic diacritic marks (harakat) to narration text before TTS synthesis, improving pronunciation quality. Tashkeel only affects TTS input; subtitle text and `narration.txt` remain clean. (Credit: @Khaled97Sho, PR #33)
11
+ - **Manuscript reference images** — per-segment reference images for manuscript tasks are now supported via `reference_images` + `reference_images_map` parameters, matching the same feature used in creative pipelines. (Credit: @Khaled97Sho, PR #33)
12
+ - **Language-aware narration budget** — narration length for non-CJK languages now uses a language-specific characters-per-second rate (Arabic: 4.0, CJK: 5.0, others: 12.0), fixing the "narration too short for non-CJK text" issue. (Credit: @Khaled97Sho, PR #33)
13
+ - **Configurable CORS origins** — the `AGNES_CORS_ORIGINS` environment variable (comma-separated) allows any companion tool to call the core API cross-origin. `AGNES_CORS_ENABLED` can force-enable or disable; auto-detects when origins are set.
14
+ - **China domestic API domain** — the `cn` domain in `AGNES_DOMAIN_MAP` now correctly points to `api.agnes-ai.cn` (the official China service endpoint), updated from the incorrect `apihub.agnes-ai.cn` (international-site fallback). The frontend domain picker reflects the correct hostname.
18
15
 
19
16
  ### Refactoring & Optimizations
20
17
 
21
- - **ffmpeg-first compositing chain** — the final assembly path for creative/manuscript/anchor/poetry videos was reworked from 3-4 full re-encodes into copy-concat + a single filter pass (with graceful fallback to the previous moviepy path). This is the largest performance win in the v6 line, cutting final-assembly time by roughly 3-10x on typical outputs.
22
- - **Asynchronous rate limiting with dedicated encoding thread pool** — the token bucket now offers a native async acquire path (stop-aware), and heavy encoding runs on a dedicated executor so long encoding jobs no longer starve the request path.
18
+ - **`split_manuscript_text()` extracted as a public utility** — the manuscript-splitting algorithm is now available as a standalone function, used by both the manuscript pipeline and the new preview-split endpoint.
23
19
 
24
20
  ### Bug Fixes
25
21
 
26
- - **Fixed stopping behavior** — cancelling a task no longer triggers retry backoff (up to ~2 minutes) and no longer deletes a resumable `video_id`.
27
- - **Fixed multi-Key configuration** — key IDs are now hashed from the actual key so deleting one Key from multiple configured Keys removes exactly that Key.
28
- - **Fixed event-loop freezes** — watermark re-encoding and synchronous downloads no longer block the whole service; a semaphore release bug that could permanently break the concurrency cap under low-rate-limit configurations is fixed.
29
- - **Fixed frontend issues** — a stored-XSS vector via unescaped `v-html` is closed, duplicate image-submit without guard is prevented, and fetch errors now surface readable backend messages instead of silent failures.
22
+ - **Fixed rate-limiter live-lock** — a rare synchronous acquire path race condition could stall all pipelines indefinitely. The token bucket's async path is now the primary path, with the sync path hardened to prevent the live-lock. (Credit: @Khaled97Sho, PR #33)
23
+ - **Fixed image save not persisting** — `POST /api/images/generations` now `await`s the image download, ensuring the generated image is actually written to disk before the response.
24
+ - **Fixed GitHub CodeQL alerts #43 and #47** — path-injection and information-disclosure security alerts resolved.
25
+ - **Fixed PBKDF2-HMAC-SHA256 for config key IDs** — config key IDs are now hashed with a proper key-derivation function instead of weaker hashing.
26
+ - **Fixed Docker Hub Cloudflare WAF blocking** — the Docker Hub overview update script now strips `<script>` blocks and escapes `<` in the JSON payload to bypass the WAF XSS rule.
27
+ - **Fixed China domestic site domain** — `AGNES_DOMAIN_MAP["cn"]` corrected to `api.agnes-ai.cn` per the official Agnes model catalog. (Issue #37)
30
28
 
31
29
  ---
32
30
 
33
- No configuration changes are required. Existing tasks remain resumable; task state files are unchanged in format.
31
+ **Compatibility notes**: The `core/audio/tashkeel.py` module adds a new optional dependency (`mishkal`). It is installed by default via `requirements.txt`; if missing, tashkeel silently falls back to the original text. Existing task state files are forward-compatible; no migration needed.
34
32
 
35
33
  ---
36
34
 
@@ -167,14 +167,19 @@ class AgnesRateLimiter:
167
167
  def acquire(self) -> None:
168
168
  """阻塞式获取一个令牌(同步场景 / 脚本 / 测试用)。
169
169
 
170
- 如果桶中有令牌,立即消耗并返回;否则 ``time.sleep()`` 直到令牌可用。
170
+ 如果桶中有令牌,立即消耗并返回;否则 ``time.sleep()`` 等待令牌可用。
171
+
172
+ ⚠️ 修复(回归活锁):此处采用与 ``acquire_async`` 一致的**预支语义**——
173
+ ``_try_acquire`` 已把 ``last_refill`` 预留到 ``now + wait_time`` 并清空令牌,
174
+ sleep 期间令牌尚未重新累积,再次循环调用 ``_try_acquire`` 会算出 ``elapsed≈0``
175
+ 而**永远不足 1 个令牌**,导致令牌永不补充、等待者永久卡死(多并发时逐个推挤
176
+ ``last_refill`` 还使 wait_time 越滚越大)。因此 sleep 完成后直接视为已获取。
171
177
  """
172
- while True:
173
- wait_time = self._try_acquire()
174
- if wait_time is None:
175
- return
176
- self._record_wait(wait_time)
177
- time.sleep(wait_time)
178
+ wait_time = self._try_acquire()
179
+ if wait_time is None:
180
+ return
181
+ self._record_wait(wait_time)
182
+ time.sleep(wait_time)
178
183
 
179
184
  async def acquire_async(self, stop_event: asyncio.Event | None = None) -> None:
180
185
  """异步原生获取令牌(优化路线图 2.3)。
@@ -253,6 +253,14 @@ class SubtitleSrtMixin:
253
253
  if text:
254
254
  items.append((start_s, end_s, text))
255
255
 
256
+ # PRD 1.2a:TTS 若送入了加 tashkeel 的阿拉伯语文本,cues 文本会连带
257
+ # 变音符号——字幕显示前统一剥离(仅阿拉伯语变音符号,其他语言无副作用)。
258
+ from core.audio.tashkeel import strip_diacritics
259
+
260
+ items = [(s, e, strip_diacritics(t)) for s, e, t in items]
261
+ if not items:
262
+ return ""
263
+
256
264
  if not items:
257
265
  return ""
258
266
 
@@ -594,6 +602,13 @@ class SubtitleSrtMixin:
594
602
  if not items:
595
603
  return ""
596
604
 
605
+ # PRD 1.2a:TTS 若送入了加 tashkeel 的阿拉伯语文本,cues 文本会连带
606
+ # 变音符号——字幕显示前统一剥离(仅阿拉伯语变音符号,其他语言无副作用)。
607
+ # 剥离后再做归一化对齐,保证与 plain 的 segment_texts 策略 A 字符区间匹配。
608
+ from core.audio.tashkeel import strip_diacritics
609
+
610
+ items = [(s, e, strip_diacritics(t)) for s, e, t in items]
611
+
597
612
  # 残余归一化:把最后一条 cue 的 end 钳到实际音频时长(避免尾差留白未覆盖)
598
613
  if audio_duration and audio_duration > 0 and items[-1][1] > audio_duration:
599
614
  items[-1] = (items[-1][0], audio_duration, items[-1][2])
@@ -0,0 +1,117 @@
1
+ """阿拉伯语变音符号(tashkeel/harakat)自动标注,供 TTS 更准确地朗读。
2
+
3
+ 使用 mishkal(基于规则的阿拉伯语形态分析库)而非 LLM 来生成变音符号:
4
+ LLM 方案实测不可靠——即使明确要求"只加符号、不改字母",模型仍会偶发替换
5
+ 借词拼写(如 فيديو -> ويديو)或替换同义连词(如 إذا -> إن),这对旁白配音
6
+ 是不可接受的(读出的内容会与原文不同)。mishkal 只做形态学标注,天然不会
7
+ 引入新字母,因此改用"取 mishkal 的变音符号 + 保留原文逐字符结构"的合并算法,
8
+ 以程序方式保证结果与原文在去除变音符号后完全一致。
9
+
10
+ 依赖模式:**optional-dependency**(与 json-repair 相同的"缺失自动降级"模式)。
11
+ 本模块在导入时不加载 mishkal(``_get_vocalizer`` 内延迟导入),因此即使
12
+ mishkal 未安装,``strip_diacritics`` 等公共函数仍可用;``add_tashkeel_safe``
13
+ 在 mishkal 缺失时静默回退原文。``requirements.txt`` 默认安装 mishkal 保证
14
+ 开箱即用。
15
+
16
+ Port: PR #33 by @Khaled97Sho(新增公共 ``strip_diacritics`` 供字幕隔离复用)。
17
+ """
18
+ from __future__ import annotations
19
+
20
+ import logging
21
+ import re
22
+
23
+ logger = logging.getLogger(__name__)
24
+
25
+ # 阿拉伯字母(U+0621–U+064A)
26
+ _ARABIC_LETTER_RE = re.compile(r"[ء-ي]")
27
+ # 阿拉伯语变音符号(harakat/tashkeel,U+064B–U+0670):fatha/damma/kasra 等,
28
+ # 仅标注发音、不产生语音时长。字幕与旁白导出产物必须剥离,避免污染显示。
29
+ _DIACRITIC_RE = re.compile(r"[ً-ْٰ]")
30
+
31
+ _vocalizer = None
32
+
33
+
34
+ def strip_diacritics(text: str) -> str:
35
+ """去除文本中的阿拉伯语变音符号(tashkeel/harakat)。
36
+
37
+ 用于字幕 / 旁白导出产物:加 tashkeel 的文本只应送入 TTS,字幕显示与
38
+ ``narration.txt`` 必须使用剥离变音符号的干净版本(PRD 1.2a)。
39
+ 对不含变音符号的文本(其他语言)为无副作用操作。
40
+
41
+ Args:
42
+ text: 原始文本(可为空)。
43
+
44
+ Returns:
45
+ 剥离阿拉伯语变音符号后的文本。
46
+ """
47
+ if not text:
48
+ return text
49
+ return _DIACRITIC_RE.sub("", text)
50
+
51
+
52
+ def _get_vocalizer():
53
+ global _vocalizer
54
+ if _vocalizer is None:
55
+ from mishkal.tashkeel import TashkeelClass
56
+
57
+ _vocalizer = TashkeelClass()
58
+ return _vocalizer
59
+
60
+
61
+ def _merge_tashkeel(original: str, diacritized: str) -> str:
62
+ """按原文逐字符重建结果:阿拉伯字母取自 mishkal 输出(含其变音符号),
63
+ 其余字符(标点、空格、数字等)严格保留原文,不采用 mishkal 对它们的改写。
64
+ """
65
+ d_tokens: list[tuple[str, str]] = []
66
+ i = 0
67
+ while i < len(diacritized):
68
+ ch = diacritized[i]
69
+ if _ARABIC_LETTER_RE.match(ch):
70
+ j = i + 1
71
+ diac = ""
72
+ while j < len(diacritized) and _DIACRITIC_RE.match(diacritized[j]):
73
+ diac += diacritized[j]
74
+ j += 1
75
+ d_tokens.append((ch, diac))
76
+ i = j
77
+ else:
78
+ i += 1
79
+
80
+ out = []
81
+ ti = 0
82
+ for ch in original:
83
+ if _ARABIC_LETTER_RE.match(ch):
84
+ if ti < len(d_tokens):
85
+ letter, diac = d_tokens[ti]
86
+ out.append(letter + diac)
87
+ ti += 1
88
+ else:
89
+ out.append(ch)
90
+ else:
91
+ out.append(ch)
92
+ return "".join(out)
93
+
94
+
95
+ def add_tashkeel_safe(text: str) -> str:
96
+ """为阿拉伯文文本添加变音符号,失败或校验不通过时原样返回。
97
+
98
+ 校验:合并结果去除变音符号后必须与输入(同样去除变音符号后)逐字符相等,
99
+ 否则说明 mishkal 内部出现异常输出,这种情况下放弃标注、返回原文,
100
+ 避免把错误内容送入 TTS。
101
+ """
102
+ if not text or not _ARABIC_LETTER_RE.search(text):
103
+ return text
104
+
105
+ base_text = strip_diacritics(text)
106
+ try:
107
+ vocalizer = _get_vocalizer()
108
+ raw_result = vocalizer.tashkeel(base_text)
109
+ except Exception as e:
110
+ logger.warning(f"[Tashkeel] mishkal failed, returning original text: {e}")
111
+ return text
112
+
113
+ merged = _merge_tashkeel(base_text, raw_result)
114
+ if strip_diacritics(merged) != base_text:
115
+ logger.warning("[Tashkeel] validation failed (letters changed), returning original text")
116
+ return text
117
+ return merged
@@ -180,6 +180,65 @@ def detect_text_script(text: str) -> str:
180
180
  return "unknown"
181
181
 
182
182
 
183
+ # ═══════════════════════════════════════════════════
184
+ # 语速估算(PR #33 吸收:跨脚本朗读速率差异 + 变音符号剥离)
185
+ # ═══════════════════════════════════════════════════
186
+ # 单一公共实现(PRD 1.3a):story.py / manuscript_video.py / preview_routes
187
+ # 均从本模块导入,消除此前两处 4.0/13.0 常量与脚本判断的重复副本。
188
+
189
+ # CJK 字符密度高(一字近一音节),约 4 字/秒(实测 zh-CN-Xiaoxiao ≈ 4.7,取保守值)
190
+ _CHARS_PER_SEC_CJK = 4.0
191
+ # 阿拉伯文(2026-08-31 实测校准:ar-SA-Hamed/Zariyah ≈ 10.4-10.6、ar-EG-Shakir ≈ 12.0
192
+ # 字符/秒,真实旁白含句间停顿更低,取 10.5;PR #33 原沿用的统一 13 偏快,会导致
193
+ # 旁白比画面长出约 40%,视频被迫定格等待旁白结束)
194
+ _CHARS_PER_SEC_ARABIC = 10.5
195
+ # 其余字母文字(拉丁/西里尔/泰文/天城文/孟加拉文等)统一 13 字符/秒
196
+ # (英文实测 15.7-16.8 偏快、法/西/俄未实测,13 为折中保守值,后续可分档校准)
197
+ _CHARS_PER_SEC_ALPHABETIC = 13.0
198
+
199
+
200
+ def estimate_chars_per_sec(text: str) -> float:
201
+ """按文本主要文字体系估算朗读语速(字符/秒)。
202
+
203
+ 中文/日文/韩文按 CJK 速率(4.0 字/秒);阿拉伯文按实测速率(10.5 字符/秒,
204
+ 见 ``_CHARS_PER_SEC_ARABIC`` 校准记录);其余脚本(拉丁/西里尔/泰文/
205
+ 天城文/孟加拉文等)统一按字母文字速率(13.0 字符/秒,泰文等未实测脚本
206
+ 先用统一值,后续可校准)。
207
+
208
+ Args:
209
+ text: 待估算朗读时长的文本。
210
+
211
+ Returns:
212
+ 语速(字符/秒)。
213
+ """
214
+ script = detect_text_script(text)
215
+ if script in ("zh", "ja", "ko"):
216
+ return _CHARS_PER_SEC_CJK
217
+ if script == "arabic":
218
+ return _CHARS_PER_SEC_ARABIC
219
+ return _CHARS_PER_SEC_ALPHABETIC
220
+
221
+
222
+ def duration_len(text: str) -> int:
223
+ """用于时长估算的字符数(剥离阿拉伯语变音符号后计数)。
224
+
225
+ 阿拉伯语变音符号(harakat/tashkeel)只标注发音、不产生语音时长,
226
+ 估算前必须剥离,否则加全变音符号的文本(codepoint 增加 40-60%)
227
+ 会让时长估算严重偏长,导致生成视频尾部大片静音/定格。
228
+
229
+ Args:
230
+ text: 原始文本(可为空)。
231
+
232
+ Returns:
233
+ 剥离变音符号后的字符数。
234
+ """
235
+ if not text:
236
+ return 0
237
+ from core.audio.tashkeel import strip_diacritics
238
+
239
+ return len(strip_diacritics(text))
240
+
241
+
183
242
  # edge-tts voice id 前缀 → 项目语言 code 映射。
184
243
  # 例外:Tagalog 音色在 edge-tts 中用 ISO 639-3 代码 `fil`(如 fil-PH-AngeloNeural),
185
244
  # 而项目/前端 UI 用 ISO 639-1 `tl`,故需显式归一。
package/core/config.py CHANGED
@@ -18,7 +18,7 @@ CONFIG_FILE = os.path.join(CONFIG_DIR, "config.json")
18
18
  # ═══════════════════════════════════════════════════
19
19
  # 应用版本号(v6.1 新增:发版时同步更新,见 docs/dev/release_process.md)
20
20
  # ═══════════════════════════════════════════════════
21
- APP_VERSION = "6.3.0"
21
+ APP_VERSION = "6.4.0"
22
22
 
23
23
  # 未配置 API Key 时的统一报错文案(含免费获取与在线体验兜底,全站路由共用)
24
24
  API_KEY_MISSING_MSG = (
@@ -202,6 +202,13 @@ try:
202
202
  agnes_subtitle_ass: bool = True # 2.1c 字幕 ASS 单链灰度开关
203
203
  agnes_video_poll_timeout: int = 1800 # 1.2 视频轮询总超时
204
204
 
205
+ # ── CORS(PR #33 吸收 Phase 2:可配置跨源白名单)──
206
+ # 供独立本地伴侣工具(如 agnes-simple-ui)从浏览器跨源调用本服务 API。
207
+ # agnes_cors_origins: 逗号分隔的允许源列表,空 = 不启用 CORS(默认,攻击面不变)。
208
+ # agnes_cors_enabled: None=auto(设置了 origins 才启用);显式 "false" 即使设置了 origins 也禁用。
209
+ agnes_cors_origins: str = ""
210
+ agnes_cors_enabled: bool | None = None
211
+
205
212
  # ── 运维 ──
206
213
  agnes_log_file: str = ""
207
214
  agnes_sweep_age_days: int | None = None
@@ -235,6 +242,14 @@ except ImportError: # pragma: no cover - pydantic-settings 为必备依赖,
235
242
  "0", "false", "off",
236
243
  )
237
244
  self.agnes_video_poll_timeout = int(os.environ.get("AGNES_VIDEO_POLL_TIMEOUT", "1800"))
245
+ self.agnes_cors_origins = os.environ.get("AGNES_CORS_ORIGINS", "")
246
+ _cors_enabled = os.environ.get("AGNES_CORS_ENABLED", "").strip().lower()
247
+ if _cors_enabled in ("0", "false", "off"):
248
+ self.agnes_cors_enabled = False
249
+ elif _cors_enabled in ("1", "true", "on"):
250
+ self.agnes_cors_enabled = True
251
+ else:
252
+ self.agnes_cors_enabled = None
238
253
  self.agnes_log_file = os.environ.get("AGNES_LOG_FILE", "")
239
254
  self.agnes_sweep_age_days = _env_int("AGNES_SWEEP_AGE_DAYS")
240
255
  self.agnes_config_id_hmac_key = os.environ.get(
@@ -951,9 +966,12 @@ def get_video_model_capabilities() -> dict:
951
966
  # ═══════════════════════════════════════════════════
952
967
 
953
968
  # 可用域名映射
969
+ # 注:Agnes 官方域名划分(来源:AgnesAI-Labs skills 的 model_catalog 参考值):
970
+ # - com 国际站(主):apihub.agnes-ai.com
971
+ # - cn 国内站(中国站):api.agnes-ai.cn(apihub.agnes-ai.cn 仅作国际站备用,非国内站)
954
972
  AGNES_DOMAIN_MAP = {
955
973
  "com": "https://apihub.agnes-ai.com",
956
- "cn": "https://apihub.agnes-ai.cn",
974
+ "cn": "https://api.agnes-ai.cn",
957
975
  }
958
976
 
959
977
  _DEFAULT_DOMAIN = "com"
@@ -69,13 +69,35 @@ class CreativeVideoPipeline(
69
69
  """
70
70
  super().__init__(api_key, task_id, dir_name, progress_callback, shutdown_event)
71
71
 
72
- self.screenwriter = Screenwriter(api_key=api_key, model=chat_model)
72
+ # PRD 1.4:Screenwriter language pinning——默认 en、尊重显式配置。
73
+ # 默认中文系统提示词会让模型倾向于输出中文(即使提示词写明"跟随输入语言"),
74
+ # 实测英文/阿拉伯文 idea 在中文提示词下偶发被错误写成中文故事/旁白。
75
+ # 仅当用户未显式设置 PROMPT_LANGUAGE(使用默认 zh)时固定 language="en";
76
+ # 显式配置代表用户知情选择,以用户配置为准(language=None 走 PROMPT_LANGUAGE)。
77
+ # 补充(2026-08-31 实测):英文系统提示词下模型仍会因 user prompt 内的中文章节
78
+ # 片段把非中文 idea 写成中文——流水线三处 LLM 调用(故事/脚本/旁白)额外前置
79
+ # 显式语言指令(_style_with_language_directive),与 preview 端点同机制。
80
+ from core.screenwriter import is_prompt_language_explicit
81
+
82
+ _sw_language = None if is_prompt_language_explicit() else "en"
83
+ self.screenwriter = Screenwriter(api_key=api_key, model=chat_model, language=_sw_language)
73
84
  self.image_generator = AgnesImageAPI(api_key=api_key, model=image_model)
74
85
  self.video_generator = AgnesVideoAPI(api_key=api_key, model=video_model)
75
86
  self.video_generator.shutdown_event = shutdown_event
76
87
 
77
88
  self._state: Optional[CreativeVideoTask] = None
78
89
 
90
+ def _style_with_language_directive(self) -> str:
91
+ """用户 style 前置基于 idea 文字体系的显式语言指令。
92
+
93
+ idea 为非中文脚本时返回「指令 + style」,中文 idea 原样返回 style
94
+ (指令为空串)。指令确保故事/旁白跟随输入语言,同时保留"场景视觉提示词
95
+ 用英文"的既有约定。
96
+ """
97
+ from core.screenwriter import build_input_language_directive
98
+
99
+ return build_input_language_directive(self._state.idea) + self._state.style
100
+
79
101
  # ------------------------------------------------------------------
80
102
  # Properties
81
103
  # ------------------------------------------------------------------
@@ -6,6 +6,7 @@ import os
6
6
  import re
7
7
  from typing import Optional
8
8
 
9
+ from core.audio.voices import estimate_chars_per_sec
9
10
  from core.compositor.concatenator import VideoConcatenator
10
11
  from core.screenwriter import clean_narration_text
11
12
  from models.task import StepStatus
@@ -13,7 +14,6 @@ from models.task import StepStatus
13
14
  logger = logging.getLogger(__name__)
14
15
 
15
16
 
16
- _CHARS_PER_SEC = 4.0
17
17
  _SENTENCE_BOUNDARY_RE = re.compile(r"(?<=[。!?.!?])")
18
18
 
19
19
  # 音频/字幕阶段起始进度(阶段内线性插值)
@@ -120,15 +120,16 @@ class AudioStepsMixin:
120
120
  # On resume, re-trim existing narrations (old untrimmed data may persist)
121
121
  if self._state.narrations:
122
122
  _scenes = self._state.scenes
123
+ _cps = estimate_chars_per_sec(story)
123
124
  needs_update = any(
124
- len(n) > max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _CHARS_PER_SEC), 20)
125
+ len(n) > max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _cps), 20)
125
126
  for i, n in enumerate(self._state.narrations)
126
127
  )
127
128
  if needs_update:
128
129
  self._state.narrations = [
129
130
  _trim_to_sentence(
130
131
  n,
131
- max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _CHARS_PER_SEC), 20),
132
+ max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _cps), 20),
132
133
  )
133
134
  for i, n in enumerate(self._state.narrations)
134
135
  ]
@@ -155,12 +156,13 @@ class AudioStepsMixin:
155
156
  narrations.append("\n".join(narrative_paras[idx : idx + count]))
156
157
  idx += count
157
158
 
158
- # Trim each narration to fit within its scene's duration * 4 chars/sec speaking rate
159
+ # Trim each narration to fit within its scene's duration * estimated speaking rate
159
160
  _scenes = self._state.scenes
161
+ _cps = estimate_chars_per_sec(story)
160
162
  narrations = [
161
163
  _trim_to_sentence(
162
164
  n,
163
- max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _CHARS_PER_SEC), 20),
165
+ max(int((_scenes[i].duration if i < len(_scenes) else self._state.video_duration) * _cps), 20),
164
166
  )
165
167
  for i, n in enumerate(narrations)
166
168
  ]
@@ -202,12 +204,12 @@ class AudioStepsMixin:
202
204
  story,
203
205
  scenes,
204
206
  total_duration,
205
- self._state.style,
207
+ self._style_with_language_directive(),
206
208
  )
207
209
 
208
210
  if not narration or len(narration) < 5:
209
211
  logger.warning("[Pipeline] LLM returned empty/short narration, using cleaned story fallback")
210
- max_chars = max(int(total_duration * _CHARS_PER_SEC), 40)
212
+ max_chars = max(int(total_duration * estimate_chars_per_sec(story)), 40)
211
213
  narration = clean_narration_text(_trim_to_sentence(story, max_chars)) if story else ""
212
214
 
213
215
  self._state.narrations = [narration]
@@ -257,13 +259,28 @@ class AudioStepsMixin:
257
259
  # v5.x 产物规范前置:导出旁白纯文本(供外部 Agent/工具处理)
258
260
  narration_text = self._state.narrations[0] if self._state.narrations else ""
259
261
  self._save_narration_txt(narration_text, combined_audio)
262
+
263
+ # PRD 1.2a:tashkeel 只作用于送入 TTS 的文本,字幕/旁白导出产物用 plain 版本。
264
+ # state.narrations 始终保留无变音符号文本(字幕与 narration.txt 由此生成),
265
+ # 仅 narration_tts 加 tashkeel 送 TTS(含续传重采 cues)。
266
+ narration_tts = narration_text
267
+ if self._state.audio_config.add_tashkeel:
268
+ from core.audio.tashkeel import add_tashkeel_safe
269
+
270
+ narration_tts = add_tashkeel_safe(narration_text)
271
+ if narration_tts != narration_text:
272
+ logger.info(
273
+ "[Pipeline] Step audio: tashkeel applied "
274
+ "(plain kept for subtitles/export, %d chars)", len(narration_text),
275
+ )
276
+
260
277
  if os.path.exists(combined_audio) and os.path.getsize(combined_audio) > 0:
261
278
  logger.info("[Pipeline] Step audio: SKIP (file exists)")
262
279
  self._state.step_audio = StepStatus.COMPLETED
263
280
  self.task_manager.update_step("step_audio", StepStatus.COMPLETED)
264
281
  # 续传:音频已存在则仅重采 cues,避免字幕退回 legacy 启发式
265
282
  return await self._recover_sub_maker(
266
- narration_text,
283
+ narration_tts,
267
284
  self._state.audio_config, self._state.subtitle_config,
268
285
  combined_audio,
269
286
  )
@@ -283,7 +300,7 @@ class AudioStepsMixin:
283
300
 
284
301
  sub_maker = await self._generate_audio_with_fallback(
285
302
  output_path=combined_audio,
286
- text=narration_text,
303
+ text=narration_tts,
287
304
  audio_config=self._state.audio_config,
288
305
  subtitle_config=self._state.subtitle_config,
289
306
  duration_sec=total_duration,
@@ -254,7 +254,7 @@ class ScriptStepsMixin:
254
254
  self.screenwriter.develop_story,
255
255
  self._state.idea,
256
256
  "",
257
- self._state.style,
257
+ self._style_with_language_directive(),
258
258
  image_context,
259
259
  self._state.scene_count,
260
260
  self._state.scene_durations,
@@ -377,7 +377,7 @@ class ScriptStepsMixin:
377
377
  await self._emit("script", "running", "正在编写脚本...", _PROGRESS_SCRIPT_START)
378
378
  scenes = await asyncio.to_thread(
379
379
  self.screenwriter.write_script, story, "",
380
- self._state.style,
380
+ self._style_with_language_directive(),
381
381
  self._state.scene_count,
382
382
  self._state.scene_durations,
383
383
  )