@danhachuel/thunderbolt 0.3.76 → 0.3.77
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/MANUAL-INSTALACAO.md +11 -5
- package/README.md +7 -1
- package/app/main.py +9 -0
- package/hermes_ui/pipeline_worker.py +31 -3
- package/integrations/moneyprinter_config.py +25 -1
- package/package.json +1 -1
- package/seed/skills/azure_tts_chunked.py +88 -31
- package/seed/skills/mpt_agent.py +107 -24
package/MANUAL-INSTALACAO.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Este manual descreve a instalação local da UI Thunderbolt, baseada no MoneyPrinterTurbo, utilizando o pacote npm `@danhachuel/thunderbolt`. O fluxo recomendado instala automaticamente o ambiente Python, as dependências da aplicação, as dependências do MoneyPrinterTurbo, o Streamlit e o suporte FFmpeg através de `imageio-ffmpeg`.
|
|
4
4
|
|
|
5
|
-
> **Versão deste manual:** 0.3.
|
|
5
|
+
> **Versão deste manual:** 0.3.77
|
|
6
6
|
> **Pacote npm:** `@danhachuel/thunderbolt`
|
|
7
7
|
> **Porta padrão da UI:** `localhost:3030`
|
|
8
8
|
> **Repositório:** [github.com/DanHachuel/thunderbolt](https://github.com/DanHachuel/thunderbolt)
|
|
@@ -132,13 +132,13 @@ Execute:
|
|
|
132
132
|
Windows PowerShell ou MobaXterm:
|
|
133
133
|
|
|
134
134
|
```powershell
|
|
135
|
-
npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.
|
|
135
|
+
npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.77 install
|
|
136
136
|
```
|
|
137
137
|
|
|
138
138
|
Linux/macOS:
|
|
139
139
|
|
|
140
140
|
```bash
|
|
141
|
-
npx --yes --prefer-online @danhachuel/thunderbolt@0.3.
|
|
141
|
+
npx --yes --prefer-online @danhachuel/thunderbolt@0.3.77 install
|
|
142
142
|
```
|
|
143
143
|
|
|
144
144
|
A instalação normal é **segura para actualizações**: preserva `storage`, Blueprints, Brandings, configurações e artefactos do utilizador. Remove apenas `.venv`, o clone técnico do MoneyPrinterTurbo e dependências que serão recriadas. Uma pasta antiga sem dados do utilizador, como `C:\Users\<utilizador>\AppData\Local\hermes` da tentativa incompleta, pode ser removida; uma pasta antiga que contenha Blueprints, Brandings ou storage é preservada e apenas avisada no terminal. Feche processos Python, Node, Streamlit e MobaXterm que estejam a usar as pastas antes de executar.
|
|
@@ -146,7 +146,7 @@ A instalação normal é **segura para actualizações**: preserva `storage`, Bl
|
|
|
146
146
|
Se quiser apagar absolutamente tudo de forma intencional, use o comando destrutivo separado:
|
|
147
147
|
|
|
148
148
|
```powershell
|
|
149
|
-
npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.
|
|
149
|
+
npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.77 install --purge-data
|
|
150
150
|
```
|
|
151
151
|
|
|
152
152
|
O parâmetro `--purge-data` apaga Blueprints, Brandings, configurações, storage e artefactos locais. Não o use numa actualização normal.
|
|
@@ -574,7 +574,7 @@ Ao abrir a página, o Thunderbolt não prepara dados públicos, não descarrega
|
|
|
574
574
|
|
|
575
575
|
Os parâmetros da UI são número de clusters entre 2 e 10, suporte mínimo entre 0,01 e 0,50, país, engagement, intervalo de datas e tags, todos dentro da área principal da aba. O núcleo normaliza os dados, calcula engagement, aplica filtros, faz transformação logarítmica e standardização, executa K-Means e calcula itemsets/regras com FP-Growth. Não são apresentados resultados até ao primeiro clique em **Analisar Nichos**; o mesmo botão aplica alterações posteriores aos filtros. Os resultados são DataFrames de clusters, itemsets frequentes, regras de associação e dados analisados; o gráfico de dispersão é criado nativamente com Plotly.
|
|
576
576
|
|
|
577
|
-
As dependências adicionais — `scikit-learn`, `mlxtend`, `plotly`, `seaborn`, `matplotlib` e `kagglehub` — são instaladas pelo procedimento normal de `npx`. Em instalações existentes, execute novamente `npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.
|
|
577
|
+
As dependências adicionais — `scikit-learn`, `mlxtend`, `plotly`, `seaborn`, `matplotlib` e `kagglehub` — são instaladas pelo procedimento normal de `npx`. Em instalações existentes, execute novamente `npx.cmd --yes --prefer-online @danhachuel/thunderbolt@0.3.77 install`; o instalador detecta e reutiliza o que já estiver válido.
|
|
578
578
|
|
|
579
579
|
### Niche Finder Apify
|
|
580
580
|
|
|
@@ -811,3 +811,9 @@ O worker local guarda a etapa actual e os artefactos de cada tarefa. Ao reinicia
|
|
|
811
811
|
|
|
812
812
|
### Selector de modelos LLM
|
|
813
813
|
No card **OpenAI / NVIDIA NIM**, o campo **Modelo** é uma lista suspensa. Use **Consultar modelos** para actualizar as opções do endpoint; se o identificador não estiver disponível, escolha **Escrever modelo manualmente**.
|
|
814
|
+
|
|
815
|
+
## 16. Diagnóstico prioritário do Azure Speech V2
|
|
816
|
+
|
|
817
|
+
Quando Azure Speech SDK V2 estiver seleccionado, o Thunderbolt prepara a narração em segmentos antes de iniciar o MoneyPrinterTurbo. O tamanho interno dos segmentos é reduzido automaticamente para velocidades de fala lentas, o áudio final é validado e a CLI recebe `--custom-audio-file`. A geração é bloqueada se não houver confirmação dos segmentos; isto evita regressar à chamada monolítica do Azure que pode falhar com `maximum media duration of 600000ms`.
|
|
818
|
+
|
|
819
|
+
Na interface, uma tarefa em execução mostra a última actividade do helper e o tempo decorrido da etapa Vídeo. Se ocorrer uma falha, consulte o caminho `LOG_FILE` ou `RESULT_FILE` apresentado no erro. A linha `updated configuration fields` é apenas sincronização inicial de configuração e não é a causa terminal. Nunca copie chaves, cookies, tokens ou o conteúdo integral do roteiro para pedidos de suporte.
|
package/README.md
CHANGED
|
@@ -107,7 +107,7 @@ No topo da área principal da aplicação existe o menu nativo de idioma no padr
|
|
|
107
107
|
|
|
108
108
|
A UI suporta os temas **Dark** e **Light** através do menu nativo de três pontos do Streamlit, no local original do toolbar. Não existe um selector Theme adicional dentro da página. A configuração distribuída em `.streamlit/config.toml` disponibiliza as variantes nomeadas **Dark** e **Light**, e o menu nativo continua responsável por alternar entre os modos, seguindo o padrão do [MoneyPrinterTurbo](https://github.com/harry0703/MoneyPrinterTurbo). O CSS próprio do Thunderbolt usa cores semânticas, `currentColor` e `color-mix` para acompanhar o tema activo, sem alterar a posição nem a funcionalidade do toolbar, do botão Deploy e do menu principal.
|
|
109
109
|
|
|
110
|
-
## Navegação da UI 0.3.
|
|
110
|
+
## Navegação da UI 0.3.77
|
|
111
111
|
|
|
112
112
|
A barra lateral apresenta os níveis principais nesta ordem: **Início**, **Automação**, **Niche Finder**, **Canais/Perfis (Vídeos)**, **Pipeline Vídeos**, **Pipeline Música**, **AI Influencers**, **Edição**, **Growth**, **Documentação** e **Configurações**. **Canais/Perfis (Vídeos)** é expansível e contém **Canais YouTube**, **Blueprints Youtube**, **Contas TikTok**, **Prompt Masters** e **Facebook Pages**, nessa ordem. O menu **Arquivos Base** foi removido por ficar vazio.
|
|
113
113
|
**Pipeline Vídeos** contém **Criação de Vídeos**, **Backlog Vídeos**, **Roteiros**, **Thumbnails** e **Upload**; **Pipeline Música** contém **Criação de Músicas** e **Upload Música**; **Automação** contém **Automação Youtube**; **Niche Finder** contém **Niche Finder Kaggle** e **Niche Finder Apify**; **AI Influencers** contém **Personagens**, **Geração de Conteúdo IA**, **Motion Control**, **UGC Products** e **Redes Sociais**. O Início reúne o dashboard e as filas do Pipeline, sem botões de acções rápidas.
|
|
@@ -426,3 +426,9 @@ A pipeline do worker usa agora um orquestrador local em cascata, com artefactos
|
|
|
426
426
|
|
|
427
427
|
## Selector de modelos LLM
|
|
428
428
|
O campo **Modelo** em **Configuração API > API Keys > LLM — providers e modelos** é apresentado como lista suspensa permanente. A lista usa os modelos descobertos pelo endpoint e preserva o modelo guardado; a opção manual continua disponível apenas como fallback explícito.
|
|
429
|
+
|
|
430
|
+
## Correcção prioritária de geração de vídeo
|
|
431
|
+
|
|
432
|
+
No fluxo MoneyPrinterTurbo, a voz Azure Speech SDK V2 é preparada antes da CLI upstream. O Thunderbolt divide o roteiro em segmentos conservadores ajustados à velocidade, concatena o áudio localmente e só inicia o motor depois de confirmar o ficheiro e a contagem de segmentos. Se essa preparação falhar, a tarefa é interrompida e não existe fallback silencioso para uma chamada monolítica sujeita ao limite de 600000 ms.
|
|
433
|
+
|
|
434
|
+
Durante a execução, o backlog mostra a actividade recente do helper e o tempo da etapa Vídeo. Em caso de falha, a mensagem terminal é extraída do output accionável, enquanto as linhas `using existing project` e `updated configuration fields` deixam de ser tratadas como causa. Os caminhos do log e do manifesto são preservados mesmo quando a falha ocorre antes da criação do MP4.
|
package/app/main.py
CHANGED
|
@@ -3582,6 +3582,15 @@ def _render_video_task_state(task: dict[str, Any]) -> None:
|
|
|
3582
3582
|
st.write(state or "—")
|
|
3583
3583
|
st.caption(VIDEO_TASK_STATE_LABELS.get(state, state.replace("_", " ").capitalize() or "Desconhecido"))
|
|
3584
3584
|
st.progress(progress, text=f"{progress}%")
|
|
3585
|
+
helper_status = str(task.get("video_helper_status") or "").strip()
|
|
3586
|
+
if state == "doing" and helper_status:
|
|
3587
|
+
st.caption(f"Actividade: {helper_status[-240:]}")
|
|
3588
|
+
if state == "doing" and task.get("video_elapsed_seconds") is not None:
|
|
3589
|
+
try:
|
|
3590
|
+
elapsed_seconds = max(0, int(task.get("video_elapsed_seconds") or 0))
|
|
3591
|
+
st.caption(f"Tempo da etapa Vídeo: {elapsed_seconds // 60}m {elapsed_seconds % 60:02d}s")
|
|
3592
|
+
except (TypeError, ValueError):
|
|
3593
|
+
pass
|
|
3585
3594
|
if task.get("error"):
|
|
3586
3595
|
st.caption(str(task.get("error"))[:240])
|
|
3587
3596
|
|
|
@@ -310,12 +310,20 @@ def _helper_output_value(output: str, key: str) -> str:
|
|
|
310
310
|
|
|
311
311
|
|
|
312
312
|
def _redact_helper_output(text: str) -> str:
|
|
313
|
-
for key in ("MPT_LLM_API_KEY", "MPT_PEXELS_API_KEY", "MPT_PIXABAY_API_KEY"):
|
|
313
|
+
for key in ("MPT_LLM_API_KEY", "MPT_PEXELS_API_KEY", "MPT_PIXABAY_API_KEY", "AZURE_SPEECH_KEY", "AZURE_SPEECH_REGION"):
|
|
314
314
|
secret = os.environ.get(key, "").strip()
|
|
315
315
|
if secret:
|
|
316
316
|
text = text.replace(secret, "[redacted]")
|
|
317
317
|
return text
|
|
318
318
|
|
|
319
|
+
def _terminal_helper_detail(output: str) -> str:
|
|
320
|
+
"""Return the last actionable helper lines, not startup/configuration noise."""
|
|
321
|
+
lines = [_redact_helper_output(line).strip() for line in str(output or "").splitlines()]
|
|
322
|
+
lines = [line for line in lines if line]
|
|
323
|
+
noise_markers = ("using existing project:", "updated configuration fields:", "starting video generation, task id:")
|
|
324
|
+
actionable = [line for line in lines if not any(marker in line.casefold() for marker in noise_markers)]
|
|
325
|
+
return "\n".join((actionable or lines)[-12:]).strip()
|
|
326
|
+
|
|
319
327
|
|
|
320
328
|
def _helper_failure_markers(output: str) -> dict[str, Any]:
|
|
321
329
|
missing = [item.strip() for item in re.findall(r"(?m)^MISSING=(.+)$", output) if item.strip()]
|
|
@@ -378,6 +386,16 @@ def _failure_attribution(
|
|
|
378
386
|
"failure_stage": stage,
|
|
379
387
|
}
|
|
380
388
|
|
|
389
|
+
if stage == "video" and any(marker in combined for marker in ("azure speech sdk v2", "azure speech v2", "azure_tts_v2")):
|
|
390
|
+
return {
|
|
391
|
+
"failure_api": "Azure Speech SDK V2 API",
|
|
392
|
+
"failure_provider": "azure_speech",
|
|
393
|
+
"failure_service": "Narração TTS — segmentação Azure",
|
|
394
|
+
"failure_route": route,
|
|
395
|
+
"failure_config_fields": "",
|
|
396
|
+
"failure_stage": stage,
|
|
397
|
+
}
|
|
398
|
+
|
|
381
399
|
if stage == "video" and any(marker in combined for marker in ("edge_tts", "edge tts", "azure_tts_v1", "azure speech")):
|
|
382
400
|
return {
|
|
383
401
|
"failure_api": "Azure Speech / edge_tts API",
|
|
@@ -707,10 +725,13 @@ def _moneyprinter_cli_args(task: dict[str, Any], route: str, settings: dict[str,
|
|
|
707
725
|
args.append("--match-materials-to-script")
|
|
708
726
|
|
|
709
727
|
voice_mode = str(generation_settings.get("voiceover_mode") or "").strip().casefold()
|
|
728
|
+
azure_service = str(generation_settings.get("voiceover_service") or "").strip().casefold()
|
|
710
729
|
voice = str(task.get("voice") or generation_settings.get("voice") or "").strip()
|
|
711
730
|
if voice_mode == "none" or voice_mode == "upload":
|
|
712
731
|
args.extend(["--voice-name", "no-voice"])
|
|
713
|
-
elif voice:
|
|
732
|
+
elif voice or azure_service in {"azure speech sdk v2", "azure speech", "azure tts v2", "azure speech sdk"}:
|
|
733
|
+
if not voice:
|
|
734
|
+
voice = "en-US-JennyNeural"
|
|
714
735
|
if _uses_azure_speech_sdk_v2(generation_settings, settings):
|
|
715
736
|
voice = _azure_speech_v2_voice_name(voice)
|
|
716
737
|
args.extend(["--voice-name", voice])
|
|
@@ -847,6 +868,7 @@ def _run_video_helper(task: dict[str, Any]) -> Path:
|
|
|
847
868
|
reader.start()
|
|
848
869
|
output_finished = False
|
|
849
870
|
last_heartbeat = 0.0
|
|
871
|
+
last_output_line = ""
|
|
850
872
|
try:
|
|
851
873
|
while True:
|
|
852
874
|
try:
|
|
@@ -854,6 +876,7 @@ def _run_video_helper(task: dict[str, Any]) -> Path:
|
|
|
854
876
|
if line is None:
|
|
855
877
|
output_finished = True
|
|
856
878
|
elif line:
|
|
879
|
+
last_output_line = _redact_helper_output(line).strip()
|
|
857
880
|
output_lines.append(line)
|
|
858
881
|
except queue.Empty:
|
|
859
882
|
pass
|
|
@@ -870,6 +893,7 @@ def _run_video_helper(task: dict[str, Any]) -> Path:
|
|
|
870
893
|
_update(
|
|
871
894
|
task_id,
|
|
872
895
|
progress=video_progress,
|
|
896
|
+
video_helper_status=last_output_line[-500:] if last_output_line else "",
|
|
873
897
|
video_elapsed_seconds=int(elapsed),
|
|
874
898
|
)
|
|
875
899
|
_worker_heartbeat(
|
|
@@ -877,6 +901,7 @@ def _run_video_helper(task: dict[str, Any]) -> Path:
|
|
|
877
901
|
status="running",
|
|
878
902
|
stage="video",
|
|
879
903
|
progress=video_progress,
|
|
904
|
+
video_helper_status=last_output_line[-500:] if last_output_line else "",
|
|
880
905
|
video_elapsed_seconds=int(elapsed),
|
|
881
906
|
)
|
|
882
907
|
last_heartbeat = elapsed
|
|
@@ -897,10 +922,13 @@ def _run_video_helper(task: dict[str, Any]) -> Path:
|
|
|
897
922
|
_persist_video_diagnostics(task, output)
|
|
898
923
|
if result_code == 10:
|
|
899
924
|
metadata = _failure_attribution(task, settings, "video", output=output)
|
|
925
|
+
detail = _terminal_helper_detail(output)
|
|
900
926
|
message = "A geração de vídeo precisa de credenciais adicionais do MoneyPrinterTurbo"
|
|
927
|
+
if detail:
|
|
928
|
+
message += f". Detalhe do helper: {detail}"
|
|
901
929
|
raise PipelineError(_failure_message(message, metadata), failure_metadata=metadata)
|
|
902
930
|
if result_code != 0:
|
|
903
|
-
detail =
|
|
931
|
+
detail = _terminal_helper_detail(output) or "erro sem detalhes devolvidos pelo helper"
|
|
904
932
|
metadata = _failure_attribution(task, settings, "video", error=detail, output=output)
|
|
905
933
|
message = f"MoneyPrinterTurbo falhou na etapa Vídeo: {detail}"
|
|
906
934
|
raise PipelineError(_failure_message(message, metadata), failure_metadata=metadata)
|
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
from __future__ import annotations
|
|
2
2
|
|
|
3
3
|
from pathlib import Path
|
|
4
|
+
import os
|
|
5
|
+
import tempfile
|
|
4
6
|
from typing import Any
|
|
5
7
|
|
|
6
8
|
from hermes_ui.languages import language_code
|
|
@@ -177,6 +179,28 @@ def build_moneyprinter_config(settings: dict[str, Any], existing: dict[str, Any]
|
|
|
177
179
|
return config
|
|
178
180
|
|
|
179
181
|
|
|
182
|
+
def _atomic_write_text(target: Path, text: str) -> None:
|
|
183
|
+
"""Replace TOML atomically so a failed write cannot truncate the prior file."""
|
|
184
|
+
temporary_path: Path | None = None
|
|
185
|
+
try:
|
|
186
|
+
with tempfile.NamedTemporaryFile(
|
|
187
|
+
mode="w", encoding="utf-8", dir=str(target.parent),
|
|
188
|
+
prefix=f".{target.name}.", suffix=".tmp", delete=False,
|
|
189
|
+
) as handle:
|
|
190
|
+
temporary_path = Path(handle.name)
|
|
191
|
+
handle.write(text)
|
|
192
|
+
handle.flush()
|
|
193
|
+
os.fsync(handle.fileno())
|
|
194
|
+
os.replace(temporary_path, target)
|
|
195
|
+
except Exception:
|
|
196
|
+
if temporary_path is not None:
|
|
197
|
+
try:
|
|
198
|
+
temporary_path.unlink(missing_ok=True)
|
|
199
|
+
except OSError:
|
|
200
|
+
pass
|
|
201
|
+
raise
|
|
202
|
+
|
|
203
|
+
|
|
180
204
|
def sync_moneyprinter_config(settings: dict[str, Any], moneyprinter_path: str) -> Path | None:
|
|
181
205
|
if not moneyprinter_path or toml is None:
|
|
182
206
|
return None
|
|
@@ -191,5 +215,5 @@ def sync_moneyprinter_config(settings: dict[str, Any], moneyprinter_path: str) -
|
|
|
191
215
|
except Exception:
|
|
192
216
|
existing = {}
|
|
193
217
|
payload = build_moneyprinter_config(settings, existing)
|
|
194
|
-
target
|
|
218
|
+
_atomic_write_text(target, toml.dumps(payload))
|
|
195
219
|
return target
|
package/package.json
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
|
-
"""Generate long Azure Speech V2 audio
|
|
1
|
+
"""Generate long Azure Speech V2 audio through conservative sequential chunks.
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
The Azure real-time TTS endpoint limits the produced audio of one request to
|
|
4
|
+
10 minutes. This helper deliberately stays far below that limit, adjusts the
|
|
5
|
+
text budget for slow voices, and concatenates the completed segments locally.
|
|
6
|
+
Secrets are read only from environment variables and are never printed.
|
|
6
7
|
"""
|
|
7
|
-
|
|
8
8
|
from __future__ import annotations
|
|
9
9
|
|
|
10
10
|
import argparse
|
|
@@ -18,12 +18,19 @@ from tempfile import TemporaryDirectory
|
|
|
18
18
|
from xml.sax.saxutils import escape
|
|
19
19
|
|
|
20
20
|
|
|
21
|
-
|
|
21
|
+
# Azure documents a 600000 ms maximum for real-time TTS. This is an internal
|
|
22
|
+
# character budget, not a claim that characters map to a fixed duration. The
|
|
23
|
+
# deliberately conservative ceiling leaves room for slow voices, pauses and
|
|
24
|
+
# punctuation before the service limit can be approached.
|
|
25
|
+
MAX_CHUNK_CHARACTERS = 900
|
|
26
|
+
MIN_CHUNK_CHARACTERS = 180
|
|
22
27
|
RETRY_COUNT = 3
|
|
23
28
|
|
|
24
29
|
|
|
25
30
|
def _build_parser() -> argparse.ArgumentParser:
|
|
26
|
-
parser = argparse.ArgumentParser(
|
|
31
|
+
parser = argparse.ArgumentParser(
|
|
32
|
+
description="Synthesize Azure Speech V2 audio in conservative chunks."
|
|
33
|
+
)
|
|
27
34
|
parser.add_argument("--text-file", type=Path, required=True)
|
|
28
35
|
parser.add_argument("--output", type=Path, required=True)
|
|
29
36
|
parser.add_argument("--voice", required=True)
|
|
@@ -57,18 +64,27 @@ def _split_long_piece(piece: str, limit: int) -> list[str]:
|
|
|
57
64
|
return chunks
|
|
58
65
|
|
|
59
66
|
|
|
60
|
-
def split_text(text: str, limit: int =
|
|
61
|
-
"""Split paragraphs and sentences without sending
|
|
67
|
+
def split_text(text: str, limit: int | None = None) -> list[str]:
|
|
68
|
+
"""Split paragraphs and sentences without sending an oversized request."""
|
|
69
|
+
try:
|
|
70
|
+
effective_limit = int(limit if limit is not None else MAX_CHUNK_CHARACTERS)
|
|
71
|
+
except (TypeError, ValueError):
|
|
72
|
+
effective_limit = MAX_CHUNK_CHARACTERS
|
|
73
|
+
effective_limit = max(1, effective_limit)
|
|
62
74
|
normalized = re.sub(r"\r\n?", "\n", text).strip()
|
|
63
75
|
if not normalized:
|
|
64
76
|
return []
|
|
65
|
-
pieces = [
|
|
77
|
+
pieces = [
|
|
78
|
+
part.strip()
|
|
79
|
+
for part in re.split(r"\n+|(?<=[.!?。!?;;])\s+", normalized)
|
|
80
|
+
if part.strip()
|
|
81
|
+
]
|
|
66
82
|
chunks: list[str] = []
|
|
67
83
|
current = ""
|
|
68
84
|
for piece in pieces:
|
|
69
|
-
for fragment in _split_long_piece(piece,
|
|
85
|
+
for fragment in _split_long_piece(piece, effective_limit):
|
|
70
86
|
candidate = f"{current} {fragment}".strip()
|
|
71
|
-
if current and len(candidate) >
|
|
87
|
+
if current and len(candidate) > effective_limit:
|
|
72
88
|
chunks.append(current)
|
|
73
89
|
current = fragment
|
|
74
90
|
else:
|
|
@@ -86,8 +102,20 @@ def _normalise_rate(value: float) -> float:
|
|
|
86
102
|
return max(0.25, min(4.0, rate))
|
|
87
103
|
|
|
88
104
|
|
|
105
|
+
def chunk_character_limit(rate: float) -> int:
|
|
106
|
+
"""Return a conservative text budget adjusted to the requested speech rate."""
|
|
107
|
+
normalized_rate = _normalise_rate(rate)
|
|
108
|
+
# A slow rate stretches the audio, so reduce the text budget proportionally.
|
|
109
|
+
# A fast rate is capped: the service limit is not a reason to send huge SSML.
|
|
110
|
+
return max(
|
|
111
|
+
MIN_CHUNK_CHARACTERS,
|
|
112
|
+
min(MAX_CHUNK_CHARACTERS, int(round(MAX_CHUNK_CHARACTERS * normalized_rate))),
|
|
113
|
+
)
|
|
114
|
+
|
|
115
|
+
|
|
89
116
|
def _build_ssml(text: str, voice: str, rate: float) -> str:
|
|
90
|
-
|
|
117
|
+
parts = voice.split("-", 2)
|
|
118
|
+
locale = "-".join(parts[:2]) if len(parts) >= 2 else "en-US"
|
|
91
119
|
return (
|
|
92
120
|
'<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" '
|
|
93
121
|
f'xml:lang="{escape(locale)}">'
|
|
@@ -97,35 +125,51 @@ def _build_ssml(text: str, voice: str, rate: float) -> str:
|
|
|
97
125
|
)
|
|
98
126
|
|
|
99
127
|
|
|
100
|
-
def _synthesise_chunk(
|
|
128
|
+
def _synthesise_chunk(
|
|
129
|
+
speechsdk, text: str, voice: str, rate: float, target: Path
|
|
130
|
+
) -> None:
|
|
101
131
|
speech_key = os.environ.get("AZURE_SPEECH_KEY", "").strip()
|
|
102
132
|
region = os.environ.get("AZURE_SPEECH_REGION", "").strip()
|
|
103
133
|
if not speech_key or not region:
|
|
104
|
-
raise RuntimeError(
|
|
134
|
+
raise RuntimeError(
|
|
135
|
+
"Azure Speech SDK V2 requer AZURE_SPEECH_KEY e AZURE_SPEECH_REGION."
|
|
136
|
+
)
|
|
105
137
|
speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=region)
|
|
106
138
|
speech_config.speech_synthesis_voice_name = voice
|
|
107
139
|
speech_config.set_speech_synthesis_output_format(
|
|
108
140
|
speechsdk.SpeechSynthesisOutputFormat.Audio48Khz192KBitRateMonoMp3
|
|
109
141
|
)
|
|
110
|
-
audio_config = speechsdk.audio.AudioOutputConfig(
|
|
111
|
-
|
|
142
|
+
audio_config = speechsdk.audio.AudioOutputConfig(
|
|
143
|
+
filename=str(target), use_default_speaker=False
|
|
144
|
+
)
|
|
145
|
+
synthesizer = speechsdk.SpeechSynthesizer(
|
|
146
|
+
audio_config=audio_config, speech_config=speech_config
|
|
147
|
+
)
|
|
112
148
|
try:
|
|
113
149
|
result = synthesizer.speak_ssml_async(_build_ssml(text, voice, rate)).get()
|
|
114
150
|
finally:
|
|
115
151
|
synthesizer.close()
|
|
116
152
|
if result.reason != speechsdk.ResultReason.SynthesizingAudioCompleted:
|
|
117
|
-
details = getattr(
|
|
153
|
+
details = getattr(
|
|
154
|
+
getattr(result, "cancellation_details", None), "error_details", ""
|
|
155
|
+
)
|
|
118
156
|
reason = str(details or getattr(result, "reason", "unknown"))
|
|
119
|
-
raise RuntimeError(
|
|
157
|
+
raise RuntimeError(
|
|
158
|
+
f"Azure Speech SDK V2 não concluiu um segmento: {reason[:500]}"
|
|
159
|
+
)
|
|
120
160
|
if not target.is_file() or target.stat().st_size <= 0:
|
|
121
|
-
raise RuntimeError(
|
|
161
|
+
raise RuntimeError(
|
|
162
|
+
"Azure Speech SDK V2 terminou sem produzir o áudio do segmento."
|
|
163
|
+
)
|
|
122
164
|
|
|
123
165
|
|
|
124
166
|
def generate(text: str, voice: str, rate: float, output: Path) -> int:
|
|
125
167
|
import azure.cognitiveservices.speech as speechsdk
|
|
126
168
|
from pydub import AudioSegment
|
|
127
169
|
|
|
128
|
-
ffmpeg_binary = os.environ.get("IMAGEIO_FFMPEG_EXE", "").strip() or shutil.which(
|
|
170
|
+
ffmpeg_binary = os.environ.get("IMAGEIO_FFMPEG_EXE", "").strip() or shutil.which(
|
|
171
|
+
"ffmpeg"
|
|
172
|
+
)
|
|
129
173
|
if not ffmpeg_binary:
|
|
130
174
|
try:
|
|
131
175
|
import imageio_ffmpeg
|
|
@@ -133,16 +177,21 @@ def generate(text: str, voice: str, rate: float, output: Path) -> int:
|
|
|
133
177
|
ffmpeg_binary = imageio_ffmpeg.get_ffmpeg_exe()
|
|
134
178
|
except Exception:
|
|
135
179
|
ffmpeg_binary = ""
|
|
136
|
-
if ffmpeg_binary:
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
180
|
+
if not ffmpeg_binary:
|
|
181
|
+
raise RuntimeError(
|
|
182
|
+
"FFmpeg não está disponível para concatenar os segmentos Azure Speech V2."
|
|
183
|
+
)
|
|
184
|
+
AudioSegment.converter = ffmpeg_binary
|
|
185
|
+
|
|
186
|
+
normalized_rate = _normalise_rate(rate)
|
|
187
|
+
limit = chunk_character_limit(normalized_rate)
|
|
188
|
+
chunks = split_text(text, limit=limit)
|
|
142
189
|
if not chunks:
|
|
143
190
|
raise RuntimeError("O roteiro não contém texto para síntese Azure Speech.")
|
|
144
191
|
output.parent.mkdir(parents=True, exist_ok=True)
|
|
145
|
-
with TemporaryDirectory(
|
|
192
|
+
with TemporaryDirectory(
|
|
193
|
+
prefix="azure-v2-chunks-", dir=str(output.parent)
|
|
194
|
+
) as temporary:
|
|
146
195
|
temporary_path = Path(temporary)
|
|
147
196
|
segment_paths: list[Path] = []
|
|
148
197
|
for index, chunk in enumerate(chunks, start=1):
|
|
@@ -150,21 +199,28 @@ def generate(text: str, voice: str, rate: float, output: Path) -> int:
|
|
|
150
199
|
last_error: Exception | None = None
|
|
151
200
|
for attempt in range(1, RETRY_COUNT + 1):
|
|
152
201
|
try:
|
|
153
|
-
|
|
202
|
+
segment_path.unlink(missing_ok=True)
|
|
203
|
+
_synthesise_chunk(
|
|
204
|
+
speechsdk, chunk, voice, normalized_rate, segment_path
|
|
205
|
+
)
|
|
154
206
|
last_error = None
|
|
155
207
|
break
|
|
156
|
-
except Exception as exc: # Azure SDK exposes provider-specific
|
|
208
|
+
except Exception as exc: # Azure SDK exposes provider-specific classes.
|
|
157
209
|
last_error = exc
|
|
158
210
|
if attempt < RETRY_COUNT:
|
|
159
211
|
time.sleep(2 ** (attempt - 1))
|
|
160
212
|
if last_error is not None:
|
|
161
|
-
raise RuntimeError(
|
|
213
|
+
raise RuntimeError(
|
|
214
|
+
f"Falha no segmento Azure Speech {index}/{len(chunks)}: {last_error}"
|
|
215
|
+
) from last_error
|
|
162
216
|
segment_paths.append(segment_path)
|
|
217
|
+
print(f"AZURE_CHUNK_PROGRESS={index}/{len(chunks)}", flush=True)
|
|
163
218
|
|
|
164
219
|
combined = AudioSegment.empty()
|
|
165
220
|
for segment_path in segment_paths:
|
|
166
221
|
combined += AudioSegment.from_file(segment_path, format="mp3")
|
|
167
222
|
combined.export(output, format="mp3", bitrate="192k")
|
|
223
|
+
|
|
168
224
|
if not output.is_file() or output.stat().st_size <= 0:
|
|
169
225
|
raise RuntimeError("A concatenação Azure Speech V2 não produziu áudio válido.")
|
|
170
226
|
return len(chunks)
|
|
@@ -180,6 +236,7 @@ def main(argv: list[str] | None = None) -> int:
|
|
|
180
236
|
return 1
|
|
181
237
|
print(f"AZURE_CHUNKED_AUDIO={args.output.resolve()}")
|
|
182
238
|
print(f"AZURE_CHUNK_COUNT={count}")
|
|
239
|
+
print(f"AZURE_CHUNK_CHARACTER_LIMIT={chunk_character_limit(args.rate)}")
|
|
183
240
|
return 0
|
|
184
241
|
|
|
185
242
|
|
package/seed/skills/mpt_agent.py
CHANGED
|
@@ -24,6 +24,7 @@ PROJECT_ARCHIVE_URL = (
|
|
|
24
24
|
"https://github.com/harry0703/MoneyPrinterTurbo/archive/refs/heads/main.zip"
|
|
25
25
|
)
|
|
26
26
|
DEFAULT_ROOT = Path.home() / "MoneyPrinterTurbo"
|
|
27
|
+
CHUNKED_SYNTHESIS_TIMEOUT_SECONDS = 15 * 60
|
|
27
28
|
DEFAULT_VOICE_NAME = "zh-CN-XiaoxiaoNeural-Female"
|
|
28
29
|
NEEDS_INPUT_EXIT_CODE = 10
|
|
29
30
|
SUPPORTED_SOURCES = {"pexels", "pixabay", "coverr", "local"}
|
|
@@ -72,6 +73,33 @@ def log(message: str) -> None:
|
|
|
72
73
|
print(f"[MoneyPrinterTurbo] {message}", flush=True)
|
|
73
74
|
|
|
74
75
|
|
|
76
|
+
def _atomic_write_text(path: Path, text: str) -> None:
|
|
77
|
+
"""Replace a text file atomically so an interrupted Windows write keeps the old file."""
|
|
78
|
+
path.parent.mkdir(parents=True, exist_ok=True)
|
|
79
|
+
temporary_path: Path | None = None
|
|
80
|
+
try:
|
|
81
|
+
with tempfile.NamedTemporaryFile(
|
|
82
|
+
mode="w",
|
|
83
|
+
encoding="utf-8",
|
|
84
|
+
dir=str(path.parent),
|
|
85
|
+
prefix=f".{path.name}.",
|
|
86
|
+
suffix=".tmp",
|
|
87
|
+
delete=False,
|
|
88
|
+
) as handle:
|
|
89
|
+
temporary_path = Path(handle.name)
|
|
90
|
+
handle.write(text)
|
|
91
|
+
handle.flush()
|
|
92
|
+
os.fsync(handle.fileno())
|
|
93
|
+
os.replace(temporary_path, path)
|
|
94
|
+
except Exception:
|
|
95
|
+
if temporary_path is not None:
|
|
96
|
+
try:
|
|
97
|
+
temporary_path.unlink(missing_ok=True)
|
|
98
|
+
except OSError:
|
|
99
|
+
pass
|
|
100
|
+
raise
|
|
101
|
+
|
|
102
|
+
|
|
75
103
|
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
|
76
104
|
parser = argparse.ArgumentParser(
|
|
77
105
|
description="Install MoneyPrinterTurbo and generate a final video from a topic."
|
|
@@ -239,7 +267,7 @@ def apply_environment_config(config_path: Path) -> None:
|
|
|
239
267
|
if pixabay_keys:
|
|
240
268
|
text = _replace_config_value(text, "pixabay_api_keys", pixabay_keys)
|
|
241
269
|
changes.append("pixabay_api_keys")
|
|
242
|
-
config_path
|
|
270
|
+
_atomic_write_text(config_path, text)
|
|
243
271
|
log("updated configuration fields: " + ", ".join(changes))
|
|
244
272
|
|
|
245
273
|
|
|
@@ -279,7 +307,7 @@ def reuse_existing_llm_provider(config_path: Path) -> str:
|
|
|
279
307
|
for provider in reusable_providers:
|
|
280
308
|
if _provider_is_ready(text, provider):
|
|
281
309
|
text = _replace_config_value(text, "llm_provider", provider)
|
|
282
|
-
config_path
|
|
310
|
+
_atomic_write_text(config_path, text)
|
|
283
311
|
log(f"reusing configured LLM provider: {provider}")
|
|
284
312
|
return provider
|
|
285
313
|
return current_provider
|
|
@@ -325,12 +353,10 @@ def _forwarded_option_value(cli_args: list[str], option: str) -> str:
|
|
|
325
353
|
def _azure_v2_voice(cli_args: list[str]) -> str:
|
|
326
354
|
"""Return the unmarked Azure voice when this invocation requests Azure V2."""
|
|
327
355
|
voice = _forwarded_option_value(cli_args, "--voice-name")
|
|
328
|
-
if "-
|
|
356
|
+
if "-v2" not in voice.casefold():
|
|
329
357
|
return ""
|
|
330
|
-
base_voice = re.sub(
|
|
331
|
-
return re.sub(r"-(
|
|
332
|
-
|
|
333
|
-
|
|
358
|
+
base_voice = re.sub("-v2", "", voice, count=1, flags=re.IGNORECASE).strip()
|
|
359
|
+
return re.sub(r"-(Female|Male)$", "", base_voice, flags=re.IGNORECASE).strip()
|
|
334
360
|
def _voice_rate(cli_args: list[str]) -> float:
|
|
335
361
|
value = _forwarded_option_value(cli_args, "--voice-rate")
|
|
336
362
|
try:
|
|
@@ -358,7 +384,7 @@ def _prepare_azure_v2_chunked_audio(
|
|
|
358
384
|
text_file = task_dir / "azure-v2-script.txt"
|
|
359
385
|
audio_file = task_dir / "azure-v2-audio.mp3"
|
|
360
386
|
text_file.parent.mkdir(parents=True, exist_ok=True)
|
|
361
|
-
text_file
|
|
387
|
+
_atomic_write_text(text_file, text.strip() + "\n")
|
|
362
388
|
chunk_script = Path(__file__).with_name("azure_tts_chunked.py")
|
|
363
389
|
command = [
|
|
364
390
|
uv,
|
|
@@ -382,24 +408,55 @@ def _prepare_azure_v2_chunked_audio(
|
|
|
382
408
|
ffmpeg_path = _toml_section_value(config_path, "app", "ffmpeg_path")
|
|
383
409
|
if ffmpeg_path:
|
|
384
410
|
environment["IMAGEIO_FFMPEG_EXE"] = ffmpeg_path
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
411
|
+
try:
|
|
412
|
+
result = subprocess.run(
|
|
413
|
+
command,
|
|
414
|
+
cwd=root,
|
|
415
|
+
env=environment,
|
|
416
|
+
stdout=subprocess.PIPE,
|
|
417
|
+
stderr=subprocess.STDOUT,
|
|
418
|
+
text=True,
|
|
419
|
+
errors="replace",
|
|
420
|
+
check=False,
|
|
421
|
+
timeout=CHUNKED_SYNTHESIS_TIMEOUT_SECONDS,
|
|
422
|
+
)
|
|
423
|
+
except subprocess.TimeoutExpired as exc:
|
|
424
|
+
raise SkillError(
|
|
425
|
+
"A síntese segmentada Azure Speech SDK V2 excedeu 15 minutos; "
|
|
426
|
+
"a chamada monolítica não será usada."
|
|
427
|
+
) from exc
|
|
428
|
+
chunk_match = re.search(r"(?m)^AZURE_CHUNK_COUNT=(\d+)$", result.stdout or "")
|
|
429
|
+
if not chunk_match:
|
|
430
|
+
detail = "\n".join((result.stdout or "").splitlines()[-12:]).strip()
|
|
431
|
+
detail = detail.replace(speech_key, "[redacted]").replace(speech_region, "[redacted]")
|
|
432
|
+
raise SkillError(
|
|
433
|
+
"Azure Speech SDK V2 não confirmou segmentos e áudio customizado. "
|
|
434
|
+
+ (detail or "O helper não devolveu detalhes.")
|
|
435
|
+
)
|
|
396
436
|
if result.returncode != 0 or not audio_file.is_file() or audio_file.stat().st_size <= 0:
|
|
397
437
|
detail = "\n".join((result.stdout or "").splitlines()[-12:]).strip()
|
|
438
|
+
detail = detail.replace(speech_key, "[redacted]").replace(speech_region, "[redacted]")
|
|
398
439
|
raise SkillError(
|
|
399
440
|
"Azure Speech SDK V2 falhou na síntese segmentada. "
|
|
400
441
|
+ (detail or "O helper não devolveu detalhes.")
|
|
401
442
|
)
|
|
402
|
-
|
|
443
|
+
chunk_count = int(chunk_match.group(1))
|
|
444
|
+
if chunk_count <= 0:
|
|
445
|
+
raise SkillError("Azure Speech SDK V2 não confirmou segmentos de áudio válidos.")
|
|
446
|
+
_atomic_write_text(
|
|
447
|
+
audio_file.with_suffix(".json"),
|
|
448
|
+
json.dumps(
|
|
449
|
+
{
|
|
450
|
+
"status": "completed",
|
|
451
|
+
"audio_file": str(audio_file.resolve()),
|
|
452
|
+
"chunk_count": chunk_count,
|
|
453
|
+
},
|
|
454
|
+
ensure_ascii=False,
|
|
455
|
+
indent=2,
|
|
456
|
+
) + "\n",
|
|
457
|
+
)
|
|
458
|
+
log(f"Azure Speech V2 segmentado: {chunk_count} segmentos; áudio customizado preparado")
|
|
459
|
+
return audio_file.resolve()
|
|
403
460
|
|
|
404
461
|
|
|
405
462
|
def missing_config(config_path: Path, cli_args: list[str]) -> tuple[str, list[str]]:
|
|
@@ -524,7 +581,7 @@ def validate_pexels_config(config_path: Path, cli_args: list[str]) -> bool:
|
|
|
524
581
|
if valid_keys:
|
|
525
582
|
if valid_keys != keys:
|
|
526
583
|
text = _replace_config_value(text, "pexels_api_keys", valid_keys)
|
|
527
|
-
config_path
|
|
584
|
+
_atomic_write_text(config_path, text)
|
|
528
585
|
log(
|
|
529
586
|
"Pexels key validation completed: "
|
|
530
587
|
f"valid={len(valid_keys)}, rejected={rejected_count}, "
|
|
@@ -621,7 +678,9 @@ def generate_video(
|
|
|
621
678
|
else ["--voice-name", DEFAULT_VOICE_NAME]
|
|
622
679
|
)
|
|
623
680
|
forwarded_args = list(cli_args)
|
|
624
|
-
|
|
681
|
+
azure_v2_requested = bool(_azure_v2_voice(forwarded_args))
|
|
682
|
+
custom_audio_requested = has_cli_option(forwarded_args, "--custom-audio-file")
|
|
683
|
+
if azure_v2_requested and not custom_audio_requested:
|
|
625
684
|
chunked_audio = _prepare_azure_v2_chunked_audio(
|
|
626
685
|
root,
|
|
627
686
|
config_path,
|
|
@@ -629,8 +688,20 @@ def generate_video(
|
|
|
629
688
|
forwarded_args,
|
|
630
689
|
uv,
|
|
631
690
|
)
|
|
632
|
-
if chunked_audio is not
|
|
633
|
-
|
|
691
|
+
if chunked_audio is None or not chunked_audio.is_file() or chunked_audio.stat().st_size <= 0:
|
|
692
|
+
raise SkillError(
|
|
693
|
+
"Azure Speech SDK V2 não produziu áudio segmentado; "
|
|
694
|
+
"a geração foi interrompida para não usar a chamada monolítica."
|
|
695
|
+
)
|
|
696
|
+
forwarded_args.extend(["--custom-audio-file", str(chunked_audio.resolve())])
|
|
697
|
+
elif azure_v2_requested and custom_audio_requested:
|
|
698
|
+
custom_audio_value = _forwarded_option_value(forwarded_args, "--custom-audio-file")
|
|
699
|
+
custom_audio_path = Path(custom_audio_value).expanduser()
|
|
700
|
+
if not custom_audio_value or not custom_audio_path.is_file() or custom_audio_path.stat().st_size <= 0:
|
|
701
|
+
raise SkillError(
|
|
702
|
+
"Azure Speech SDK V2 recebeu um áudio customizado inválido; "
|
|
703
|
+
"a geração foi interrompida para não usar a chamada monolítica."
|
|
704
|
+
)
|
|
634
705
|
command = [
|
|
635
706
|
uv,
|
|
636
707
|
"run",
|
|
@@ -750,6 +821,18 @@ def main(argv: list[str] | None = None) -> int:
|
|
|
750
821
|
root, args.subject, args.cli_args
|
|
751
822
|
)
|
|
752
823
|
except (OSError, SkillError, urllib.error.URLError, zipfile.BadZipFile) as exc:
|
|
824
|
+
manifest_path = result_manifest_path(root)
|
|
825
|
+
try:
|
|
826
|
+
manifest_payload = json.loads(manifest_path.read_text(encoding="utf-8")) if manifest_path.is_file() else {}
|
|
827
|
+
except (OSError, json.JSONDecodeError):
|
|
828
|
+
manifest_payload = {}
|
|
829
|
+
if manifest_payload.get("status") == "running":
|
|
830
|
+
manifest_payload.update({"status": "failed", "error": str(exc)[:1000]})
|
|
831
|
+
write_result_manifest(root, manifest_payload)
|
|
832
|
+
for output_key, manifest_key in (("TASK_DIR", "task_dir"), ("LOG_FILE", "log_file"), ("RESULT_FILE", "result_file")):
|
|
833
|
+
value = str(manifest_payload.get(manifest_key) or "").strip()
|
|
834
|
+
if value:
|
|
835
|
+
print(f"{output_key}={value}", file=sys.stderr)
|
|
753
836
|
print(f"MPT_ERROR={exc}", file=sys.stderr)
|
|
754
837
|
return 1
|
|
755
838
|
|