@danhachuel/thunderbolt 0.2.74 → 0.2.75

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Este manual descreve a instalação local da UI Thunderbolt, baseada no MoneyPrinterTurbo, utilizando o pacote npm `@danhachuel/thunderbolt`. O fluxo recomendado instala automaticamente o ambiente Python, as dependências da aplicação, as dependências do MoneyPrinterTurbo, o Streamlit e o suporte FFmpeg através de `imageio-ffmpeg`.
4
4
 
5
- > **Versão deste manual:** 0.2.56
5
+ > **Versão deste manual:** 0.2.75
6
6
  > **Pacote npm:** `@danhachuel/thunderbolt`
7
7
  > **Porta padrão da UI:** `localhost:3030`
8
8
  > **Repositório:** [github.com/DanHachuel/thunderbolt](https://github.com/DanHachuel/thunderbolt)
@@ -438,10 +438,12 @@ A alternativa Apify não usa o dataset, parâmetros, execução ou estado da alt
438
438
 
439
439
  O menu **AI Influencers** foi adicionado abaixo de **Edição**. As abas **Personagens**, **Redes Sociais** e **Tutorial Meta** aparecem nessa ordem. **Personagens** e **Redes Sociais** mostram apenas uma mensagem de reserva para desenvolvimento futuro; **Tutorial Meta** apresenta o guia local de configuração de Instagram e credenciais Meta para automações com n8n.
440
440
 
441
- ## Edição: Limpador de Metadados, Cortes e Editor Python
441
+ ## Edição: Limpador de Metadados, Cortes, Editor Python e Download Mídia
442
442
 
443
443
  A aba **Limpador de Metadados** continua funcional e foi movida para **Edição**. A aba **Cortes** é um Clip Generator local inspirado no [OpenShorts](https://github.com/mutonby/openshorts): permite upload de vídeo, URL directa, vídeos gerados ou pasta local, formatos 9:16/1:1/16:9, opções avançadas, modo manual ou automático por segmentos locais e confirmação de direitos antes da geração. Os clips são processados com FFmpeg, guardados em `storage/cuts/runs/<id>/`, apresentados com preview e downloads individual/ZIP, acompanhados de manifesto JSON e histórico. A aba **Editor Python** é funcional e permite escolher vídeos gerados, indicar uma pasta local ou fazer upload manual; as operações são manuais e criam cópias sem alterar os originais.
444
444
 
445
+ A aba **Download Mídia** usa a API Python do [yt-dlp](https://github.com/yt-dlp/yt-dlp) para descarregar vídeos ou áudio a partir de URLs públicas. Aceita uma URL por linha, permite escolher qualidade/contentor, formato de áudio, legendas, metadados e processamento de playlists, que fica desactivado por padrão. Os resultados são guardados em `storage/downloads/`, o histórico em `storage/state/media_downloads.json` e o progresso é apresentado durante a operação. A combinação de streams e a conversão de áudio podem exigir FFmpeg. A ferramenta não aceita cookies, tokens ou opções de linha de comandos introduzidas pelo utilizador.
446
+
445
447
  ## Editor Python baseado no PYEdit
446
448
 
447
449
  O **Editor Python** adapta o recorte do [PYEdit](https://github.com/Congren/PYEdit) ao Thunderbolt. Na subaba **Vídeos**, pode seleccionar um vídeo já gerado e registado nos artefactos da pipeline, indicar uma pasta local de vídeos ou fazer upload manual. As operações disponíveis são cortar trecho, remover áudio, extrair áudio, substituir áudio, alterar velocidade e redimensionar vídeo.
package/README.md CHANGED
@@ -38,14 +38,16 @@ Os dados são segredos de sessão. Os valores não aparecem em tabelas ou logs,
38
38
 
39
39
  Os adaptadores do MoneyPrinterTurbo e de publicação nas plataformas são ligados pelas configurações locais e pelos pontos de integração em `integrations/`. A UI não inventa dados quando um serviço externo ou credencial não está disponível.
40
40
 
41
- ## Navegação da UI 0.2.72
41
+ ## Navegação da UI 0.2.75
42
42
 
43
- A barra lateral mantém os níveis principais, nesta ordem: **Início**, **Niche Finder**, **Pipeline**, **Pipeline TikTok**, **Automação**, **Edição**, **AI Influencers** e **Configurações**. **Pipeline** é expansível e contém **Criação de Vídeos**, **Criação de Músicas**, **Roteiros** e **Upload**. **Automação** também é expansível e contém **Automação Youtube**. **Edição** é expansível e contém **Limpador de Metadados**, **Cortes** e **Editor Python**, nessa ordem. **AI Influencers** é expansível e contém **Personagens**, **Redes Sociais** e **Tutorial Meta**, nessa ordem. **Niche Finder** é expansível e contém **Niche Finder Kaggle** e **Niche Finder Apify**. **Configurações** é expansível e contém **Canais Youtube**, **Blueprints Youtube**, **MCP**, **Contas Google**, **Configuração API** e **Notificações**. O Início reúne o dashboard e as filas do Pipeline, sem botões de acções rápidas.
43
+ A barra lateral mantém os níveis principais, nesta ordem: **Início**, **Niche Finder**, **Pipeline**, **Pipeline TikTok**, **Automação**, **Edição**, **AI Influencers** e **Configurações**. **Pipeline** é expansível e contém **Criação de Vídeos**, **Criação de Músicas**, **Roteiros** e **Upload**. **Automação** também é expansível e contém **Automação Youtube**. **Edição** é expansível e contém **Limpador de Metadados**, **Cortes**, **Editor Python** e **Download Mídia**, nessa ordem. **AI Influencers** é expansível e contém **Personagens**, **Redes Sociais** e **Tutorial Meta**, nessa ordem. **Niche Finder** é expansível e contém **Niche Finder Kaggle** e **Niche Finder Apify**. **Configurações** é expansível e contém **Canais Youtube**, **Blueprints Youtube**, **MCP**, **Contas Google**, **Configuração API** e **Notificações**. O Início reúne o dashboard e as filas do Pipeline, sem botões de acções rápidas.
44
44
 
45
45
  Dentro de **Configurações > Contas Google**, a UI contém os cartões expansíveis de contas Google/YouTube, `sessionInfo`, documentos de Upload directo, `INNERTUBE_API_KEY`, o formulário **Adicionar outra conta Gmail** e a configuração global do YouTube (OAuth Client ID, OAuth Client Secret e YouTube Data API Key). A página **Configuração API** contém as restantes API Keys, providers, modelos, serviços, materiais, Nano Banana, TikTok, Postiz e o **Teste de vozes**.
46
46
 
47
47
  A página **AI Influencers > Tutorial Meta** apresenta o guia de configuração de uma conta Instagram profissional e das credenciais Meta para automações com n8n, distribuído localmente em `seed/references/guide-instagram.md` e com ligação para a [fonte original no GitHub](https://github.com/gyoridavid/ai_agents_az/blob/main/episode_8/guide-instagram.md). A página **Configurações > Notificações** mantém um histórico persistente de conclusões e falhas, reconcilia estados escritos por componentes locais e disponibiliza um checkbox independente para cada operação mapeada.
48
48
 
49
+ A página **Edição > Download Mídia** utiliza a API Python do [yt-dlp](https://github.com/yt-dlp/yt-dlp) para descarregar vídeos e áudio de URLs públicas, com qualidade, contentor, formato de áudio, legendas, metadados, playlists, progresso e histórico local em `storage/downloads/` e `storage/state/media_downloads.json`. Conversão e combinação de streams podem exigir FFmpeg.
50
+
49
51
  ## Canais Youtube — edição por cartão e vídeos recentes
50
52
 
51
53
  A página **Canais Youtube** mantém o cadastro e a importação existentes, mas cada cartão agora tem o botão **Editar**. O editor permite alterar nome, URL, handle, idioma, estilo wide, **Canais de Referência / Nicho**, **Prompts do Canal** (Blueprint padrão), **Narrador** (voz padrão), conta Google do Upload directo, descrição e Automação ON/horário. O nicho aparece imediatamente abaixo do nome do canal no cartão; quando não existe, a UI mostra **SEM NICHO CONFIGURADO**.
@@ -9,3 +9,7 @@ Thunderbolt adapts the data-analysis ideas and parts of the clustering/tag-assoc
9
9
  > The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
10
10
  >
11
11
  > THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
12
+
13
+ ## yt-dlp
14
+
15
+ Thunderbolt depends on and embeds the Python API of [yt-dlp](https://github.com/yt-dlp/yt-dlp) for public media downloads. yt-dlp is distributed under [The Unlicense](https://github.com/yt-dlp/yt-dlp/blob/master/LICENSE), subject to the notices and licensing information maintained by the upstream project. This dependency is separate from Thunderbolt's MIT License.
package/app/main.py CHANGED
@@ -22,7 +22,7 @@ except (OSError, json.JSONDecodeError):
22
22
 
23
23
  from hermes_ui.domain import STAGES, create_batch, create_channel, create_tasks_for_batch, delete_channel, pipeline_summary, set_channel_defaults, transition_task, update_channel, update_channel_video
24
24
  from hermes_ui.automation_worker import load_worker_status
25
- from hermes_ui.storage import BLUEPRINTS, STORAGE, TIKTOK_PROMPT_MASTERS, ensure_storage, get_display_name, list_blueprint_files, list_prompt_master_files, load_blueprint_file, load_prompt_master_file, now, read_json, set_display_name, write_json
25
+ from hermes_ui.storage import BLUEPRINTS, MEDIA_DOWNLOADS, STORAGE, TIKTOK_PROMPT_MASTERS, ensure_storage, get_display_name, list_blueprint_files, list_prompt_master_files, load_blueprint_file, load_prompt_master_file, now, read_json, set_display_name, write_json
26
26
  from app.modules.niche_finder.apify import ApifyError, DEFAULT_ACTOR_ID, abort_actor_run, build_actor_input, get_dataset_items, normalize_video_items, start_actor_run, wait_for_actor_run
27
27
  from app.modules.niche_finder.core import NicheAnalysisError, run_niche_analysis
28
28
  from app.modules.niche_finder.data_loader import DatasetError, download_kaggle_dataset
@@ -34,6 +34,7 @@ from hermes_ui.cuts import CutsError, download_direct_video_url, generate_clips,
34
34
  from hermes_ui.mcp import detect_local_service, install_skill_locally, load_integrations, load_server_config, read_packaged_skill, save_server_config, update_integration
35
35
  from hermes_ui.mcp_server import server_status, start_server, stop_server
36
36
  from hermes_ui.music import list_music_files, materialize_suno_audio, request_suno_generation, store_music_file
37
+ from hermes_ui.media_downloader import AUDIO_FORMATS, VIDEO_CONTAINERS, VIDEO_QUALITY_OPTIONS, MediaDownloadError, build_download_options, clear_media_download_history, dependency_status, download_media, list_media_downloads, media_download_file
37
38
  from hermes_ui.notifications import clear_notifications, list_notifications, mark_all_notifications_read, mark_notification_read, notification_event_catalog, notification_preferences, record_notification, reconcile_persisted_notifications, save_notification_preferences, unread_notification_count
38
39
  from hermes_ui.script_documents import list_script_documents, read_script_document, save_script_document, script_storage_path
39
40
  from hermes_ui.script_generation import generate_script_document
@@ -2110,6 +2111,120 @@ def render_edit_placeholder(page_title: str, description: str):
2110
2111
  st.info("Esta aba está reservada para desenvolvimento futuro e ainda não executa nenhuma operação.")
2111
2112
 
2112
2113
 
2114
+ def render_media_download():
2115
+ st.title("Download Mídia")
2116
+ st.caption("Baixe vídeos e áudio de URLs públicas através da API oficial do yt-dlp. Use esta ferramenta apenas com conteúdo que tem autorização para descarregar e utilizar.")
2117
+ st.markdown("Baseado em [yt-dlp](https://github.com/yt-dlp/yt-dlp), um downloader open source para vídeo e áudio.")
2118
+ dependency = dependency_status()
2119
+ if not dependency["yt_dlp"]:
2120
+ st.warning("yt-dlp não está instalado neste ambiente. Execute a instalação das dependências do Thunderbolt antes de iniciar um download.")
2121
+ st.info("A combinação de streams, conversão de áudio e incorporação de metadados pode exigir FFmpeg. Downloads longos permanecem nesta página até terminarem.")
2122
+
2123
+ with st.form("media_download_form"):
2124
+ urls_text = st.text_area("URLs para descarregar", placeholder="Uma URL http(s) por linha", height=120, key="media_download_urls")
2125
+ mode_label = st.radio("Tipo de mídia", ["Vídeo", "Áudio"], horizontal=True, key="media_download_mode")
2126
+ option_cols = st.columns(3)
2127
+ with option_cols[0]:
2128
+ if mode_label == "Vídeo":
2129
+ quality_label = st.selectbox("Qualidade", list(VIDEO_QUALITY_OPTIONS), key="media_download_quality")
2130
+ video_container = st.selectbox("Contentor", list(VIDEO_CONTAINERS), key="media_download_container")
2131
+ audio_format = "mp3"
2132
+ else:
2133
+ quality_label = "Melhor qualidade"
2134
+ video_container = "mp4"
2135
+ audio_format = st.selectbox("Formato de áudio", list(AUDIO_FORMATS), key="media_download_audio_format")
2136
+ with option_cols[1]:
2137
+ allow_playlist = st.checkbox("Permitir playlist", value=False, key="media_download_allow_playlist")
2138
+ download_subtitles = st.checkbox("Descarregar legendas", value=False, key="media_download_subtitles")
2139
+ with option_cols[2]:
2140
+ embed_metadata = st.checkbox("Incorporar metadados", value=False, key="media_download_embed_metadata")
2141
+ st.caption("Playlist desactivada por padrão para evitar downloads acidentais em massa.")
2142
+ start_download = st.form_submit_button("Iniciar download", type="primary", use_container_width=True)
2143
+
2144
+ if start_download:
2145
+ progress = st.progress(0, text="A preparar o download…")
2146
+ progress_status = st.empty()
2147
+
2148
+ def on_progress(payload: dict[str, Any]) -> None:
2149
+ value = float(payload.get("progress") or 0)
2150
+ progress.progress(int(max(0, min(100, value))), text=f"{payload.get('status', 'processing').capitalize()} · {payload.get('current_file') or payload.get('display_url') or 'a processar'}")
2151
+ progress_status.caption(str(payload.get("hook_status") or payload.get("status") or "processing"))
2152
+
2153
+ try:
2154
+ results = download_media(
2155
+ urls_text,
2156
+ mode="video" if mode_label == "Vídeo" else "audio",
2157
+ quality=quality_label,
2158
+ container=video_container,
2159
+ audio_format=audio_format,
2160
+ allow_playlist=allow_playlist,
2161
+ download_subtitles=download_subtitles,
2162
+ embed_metadata=embed_metadata,
2163
+ progress_callback=on_progress,
2164
+ )
2165
+ st.session_state["media_download_last_results"] = results
2166
+ completed = sum(1 for item in results if item.get("status") == "completed")
2167
+ failed = len(results) - completed
2168
+ if completed:
2169
+ st.success(f"{completed} download(s) concluído(s) e guardado(s) em `{MEDIA_DOWNLOADS}`.")
2170
+ if failed:
2171
+ st.warning(f"{failed} download(s) terminou/terminaram com erro. Consulte o histórico abaixo.")
2172
+ except (MediaDownloadError, ValueError, OSError) as exc:
2173
+ progress.empty()
2174
+ st.error(str(exc))
2175
+
2176
+ latest_results = st.session_state.get("media_download_last_results", [])
2177
+ if latest_results:
2178
+ st.subheader("Resultado da última execução")
2179
+ for record in latest_results:
2180
+ with st.container(border=True):
2181
+ status_label = "Concluído" if record.get("status") == "completed" else "Falhou"
2182
+ st.write(f"**{record.get('title') or record.get('display_url') or 'Download'}** — {status_label}")
2183
+ st.caption(f"{record.get('display_url', 'URL não disponível')} · {record.get('mode', 'video')} · {record.get('completed_at') or record.get('created_at', '—')}")
2184
+ if record.get("error"):
2185
+ st.error(record["error"])
2186
+ for filename in record.get("files", []):
2187
+ output = media_download_file(record, str(filename))
2188
+ if output:
2189
+ st.download_button("Descarregar ficheiro", data=output.read_bytes(), file_name=output.name, mime="audio/*" if record.get("mode") == "audio" else "video/*", key=f"media_result_{record.get('operation_id')}_{filename}")
2190
+
2191
+ st.divider()
2192
+ st.subheader("Histórico de downloads")
2193
+ history = list_media_downloads()
2194
+ action_cols = st.columns([1, 1, 3])
2195
+ with action_cols[0]:
2196
+ if st.button("Actualizar histórico", key="media_download_refresh"):
2197
+ st.rerun()
2198
+ with action_cols[1]:
2199
+ clear_requested = st.button("Limpar histórico", key="media_download_clear")
2200
+ if clear_requested:
2201
+ st.session_state["media_download_confirm_clear"] = True
2202
+ if st.session_state.get("media_download_confirm_clear"):
2203
+ st.warning("Isto remove apenas o histórico, não os ficheiros guardados em storage/downloads.")
2204
+ confirm_cols = st.columns(2)
2205
+ with confirm_cols[0]:
2206
+ if st.button("Confirmar limpeza", type="primary", key="media_download_confirm_clear_button"):
2207
+ clear_media_download_history()
2208
+ st.session_state.pop("media_download_confirm_clear", None)
2209
+ st.rerun()
2210
+ with confirm_cols[1]:
2211
+ if st.button("Cancelar", key="media_download_cancel_clear_button"):
2212
+ st.session_state.pop("media_download_confirm_clear", None)
2213
+ st.rerun()
2214
+ if not history:
2215
+ st.caption("Ainda não existem downloads registados.")
2216
+ for record in history:
2217
+ status_label = {"completed": "Concluído", "failed": "Falhou", "processing": "Em processamento"}.get(str(record.get("status")), str(record.get("status") or "—"))
2218
+ with st.expander(f"{record.get('title') or record.get('display_url') or 'Download'} — {status_label}", expanded=False):
2219
+ st.caption(f"{record.get('display_url', 'URL não disponível')} · {record.get('mode', 'video')} · {record.get('created_at', '—')}")
2220
+ if record.get("error"):
2221
+ st.error(record["error"])
2222
+ for filename in record.get("files", []):
2223
+ output = media_download_file(record, str(filename))
2224
+ if output:
2225
+ st.download_button("Descarregar", data=output.read_bytes(), file_name=output.name, mime="audio/*" if record.get("mode") == "audio" else "video/*", key=f"media_history_{record.get('operation_id')}_{filename}")
2226
+
2227
+
2113
2228
  def render_cuts():
2114
2229
  st.title("Cortes")
2115
2230
  st.caption("Crie clips verticais, quadrados ou horizontais a partir de vídeos longos, com um fluxo local inspirado no Clip Generator do OpenShorts.")
@@ -3776,6 +3891,7 @@ def main():
3776
3891
  ("Limpador de Metadados", ":material/edit_note:", "Limpador de Metadados"),
3777
3892
  ("Cortes", ":material/content_cut:", "Cortes"),
3778
3893
  ("Editor Python", ":material/code:", "Editor Python"),
3894
+ ("Download Mídia", ":material/download:", "Download Mídia"),
3779
3895
  ]
3780
3896
  models_ai_items = [
3781
3897
  ("Personagens", ":material/person:", "Personagens"),
@@ -3882,6 +3998,7 @@ def main():
3882
3998
  "Limpador de Metadados": render_metadata_cleaner,
3883
3999
  "Cortes": render_cuts,
3884
4000
  "Editor Python": render_python_editor,
4001
+ "Download Mídia": render_media_download,
3885
4002
  "AI Influencers": lambda: render_edit_placeholder("AI Influencers", "Seleccione uma das abas AI Influencers no menu expansível."),
3886
4003
  "Tutorial Meta": render_models_ai_tutorial,
3887
4004
  "Personagens": lambda: render_edit_placeholder("Personagens", "Área reservada para a futura funcionalidade de personagens."),
@@ -0,0 +1,313 @@
1
+ from __future__ import annotations
2
+
3
+ import hashlib
4
+ import re
5
+ import uuid
6
+ from datetime import datetime, timezone
7
+ from pathlib import Path
8
+ from typing import Any, Callable, Iterable
9
+ from urllib.parse import urlsplit, urlunsplit
10
+
11
+ from . import storage
12
+ from .notifications import record_notification
13
+
14
+ try:
15
+ import yt_dlp # type: ignore
16
+ except ImportError: # pragma: no cover - exercised when the optional dependency is absent
17
+ yt_dlp = None
18
+
19
+
20
+ HISTORY_FILE = "media_downloads.json"
21
+ VIDEO_QUALITY_OPTIONS = {
22
+ "Melhor qualidade": "bv*+ba/b",
23
+ "1080p ou inferior": "bv*[height<=1080]+ba/b[height<=1080]",
24
+ "720p ou inferior": "bv*[height<=720]+ba/b[height<=720]",
25
+ "480p ou inferior": "bv*[height<=480]+ba/b[height<=480]",
26
+ }
27
+ VIDEO_CONTAINERS = ("mp4", "mkv", "webm")
28
+ AUDIO_FORMATS = ("mp3", "m4a", "wav", "opus")
29
+ ProgressCallback = Callable[[dict[str, Any]], None]
30
+
31
+
32
+ class MediaDownloadError(RuntimeError):
33
+ """Raised when a media download cannot be completed safely."""
34
+
35
+
36
+ def _now() -> str:
37
+ return datetime.now(timezone.utc).isoformat()
38
+
39
+
40
+ def _download_root() -> Path:
41
+ storage.ensure_storage()
42
+ storage.MEDIA_DOWNLOADS.mkdir(parents=True, exist_ok=True)
43
+ return storage.MEDIA_DOWNLOADS.resolve()
44
+
45
+
46
+ def _safe_download_path(value: str | Path) -> Path | None:
47
+ root = _download_root()
48
+ candidate = Path(value)
49
+ if not candidate.is_absolute():
50
+ candidate = root / candidate
51
+ try:
52
+ resolved = candidate.resolve()
53
+ resolved.relative_to(root)
54
+ except (OSError, ValueError):
55
+ return None
56
+ return resolved
57
+
58
+
59
+ def _relative_name(path: Path) -> str:
60
+ root = _download_root()
61
+ try:
62
+ return path.resolve().relative_to(root).as_posix()
63
+ except (OSError, ValueError):
64
+ return path.name
65
+
66
+
67
+ def _display_url(value: str) -> str:
68
+ parsed = urlsplit(value.strip())
69
+ if not parsed.scheme or not parsed.netloc:
70
+ return value.strip()[:120]
71
+ safe = urlunsplit((parsed.scheme, parsed.netloc, parsed.path, "", ""))
72
+ return safe[:160]
73
+
74
+
75
+ def _redact(value: Any) -> str:
76
+ text = str(value or "").strip()
77
+ text = re.sub(r"(?i)(authorization|bearer|api[_ -]?key|access[_ -]?token|cookie|session[_ -]?info)\s*[:=]?\s*[^\s,;]+", r"\1=[redacted]", text)
78
+ return text[:1200]
79
+
80
+
81
+ def normalize_urls(value: str | Iterable[str]) -> list[str]:
82
+ """Normalize one URL per line and reject unsafe/non-web inputs."""
83
+ if isinstance(value, str):
84
+ candidates = value.splitlines()
85
+ else:
86
+ candidates = list(value)
87
+ urls: list[str] = []
88
+ seen: set[str] = set()
89
+ for candidate in candidates:
90
+ url = str(candidate or "").strip()
91
+ if not url:
92
+ continue
93
+ if url.startswith("-"):
94
+ raise ValueError("Cada linha deve conter apenas uma URL; opções do yt-dlp não são aceites.")
95
+ parsed = urlsplit(url)
96
+ if parsed.scheme.lower() not in {"http", "https"} or not parsed.netloc:
97
+ raise ValueError(f"URL inválida ou não suportada: {_display_url(url)}")
98
+ normalized = urlunsplit((parsed.scheme.lower(), parsed.netloc, parsed.path, parsed.query, parsed.fragment))
99
+ if normalized not in seen:
100
+ urls.append(normalized)
101
+ seen.add(normalized)
102
+ if not urls:
103
+ raise ValueError("Introduza pelo menos uma URL http(s) para descarregar.")
104
+ return urls
105
+
106
+
107
+ def build_download_options(
108
+ *,
109
+ mode: str = "video",
110
+ quality: str = "Melhor qualidade",
111
+ container: str = "mp4",
112
+ audio_format: str = "mp3",
113
+ allow_playlist: bool = False,
114
+ download_subtitles: bool = False,
115
+ embed_metadata: bool = False,
116
+ progress_hook: Callable[[dict[str, Any]], None] | None = None,
117
+ ) -> dict[str, Any]:
118
+ """Build a constrained YoutubeDL options dictionary without user CLI flags."""
119
+ normalized_mode = str(mode or "video").strip().lower()
120
+ if normalized_mode not in {"video", "audio"}:
121
+ raise ValueError("O modo deve ser Vídeo ou Áudio.")
122
+ normalized_container = str(container or "mp4").lower()
123
+ normalized_audio = str(audio_format or "mp3").lower()
124
+ if normalized_container not in VIDEO_CONTAINERS:
125
+ raise ValueError("Contentor de vídeo não suportado.")
126
+ if normalized_audio not in AUDIO_FORMATS:
127
+ raise ValueError("Formato de áudio não suportado.")
128
+ root = _download_root()
129
+ options: dict[str, Any] = {
130
+ "outtmpl": str(root / "%(title).200B [%(id)s].%(ext)s"),
131
+ "noplaylist": not bool(allow_playlist),
132
+ "quiet": True,
133
+ "no_warnings": True,
134
+ "ignoreerrors": False,
135
+ "windowsfilenames": True,
136
+ "overwrites": False,
137
+ "paths": {"home": str(root)},
138
+ }
139
+ if progress_hook is not None:
140
+ options["progress_hooks"] = [progress_hook]
141
+ if normalized_mode == "audio":
142
+ options.update({"format": "bestaudio/best", "postprocessors": [{"key": "FFmpegExtractAudio", "preferredcodec": normalized_audio, "preferredquality": "192"}]})
143
+ else:
144
+ options.update({"format": VIDEO_QUALITY_OPTIONS.get(quality, VIDEO_QUALITY_OPTIONS["Melhor qualidade"]), "merge_output_format": normalized_container})
145
+ if download_subtitles:
146
+ options.update({"writesubtitles": True, "writeautomaticsub": True, "subtitlesformat": "best", "subtitleslangs": ["all"]})
147
+ if embed_metadata:
148
+ options["addmetadata"] = True
149
+ return options
150
+
151
+
152
+ def _read_history() -> list[dict[str, Any]]:
153
+ records = storage.read_json(HISTORY_FILE, [])
154
+ return [item for item in records if isinstance(item, dict)] if isinstance(records, list) else []
155
+
156
+
157
+ def _write_history(records: list[dict[str, Any]]) -> None:
158
+ storage.write_json(HISTORY_FILE, records[:200])
159
+
160
+
161
+ def _upsert_history(record: dict[str, Any]) -> None:
162
+ history = [item for item in _read_history() if str(item.get("operation_id")) != str(record.get("operation_id"))]
163
+ _write_history([record, *history])
164
+
165
+
166
+ def list_media_downloads(limit: int = 50) -> list[dict[str, Any]]:
167
+ return _read_history()[: max(0, int(limit))]
168
+
169
+
170
+ def clear_media_download_history() -> int:
171
+ count = len(_read_history())
172
+ _write_history([])
173
+ return count
174
+
175
+
176
+ def media_download_file(record: dict[str, Any], filename: str) -> Path | None:
177
+ """Resolve a history filename strictly inside storage/downloads."""
178
+ allowed = {str(item) for item in record.get("files", []) if item}
179
+ if filename not in allowed:
180
+ return None
181
+ path = _safe_download_path(filename)
182
+ return path if path and path.is_file() else None
183
+
184
+
185
+ def dependency_status() -> dict[str, Any]:
186
+ return {"yt_dlp": yt_dlp is not None, "ffmpeg_note": "A conversão/combinação de streams pode exigir FFmpeg."}
187
+
188
+
189
+ def _files_from_info(info: Any, started_at: float) -> list[Path]:
190
+ candidates: list[str] = []
191
+ if isinstance(info, dict):
192
+ for key in ("filepath", "_filename", "filename"):
193
+ if info.get(key):
194
+ candidates.append(str(info[key]))
195
+ requested = info.get("requested_downloads")
196
+ if isinstance(requested, list):
197
+ for item in requested:
198
+ if isinstance(item, dict):
199
+ for key in ("filepath", "_filename", "filename"):
200
+ if item.get(key):
201
+ candidates.append(str(item[key]))
202
+ entries = info.get("entries")
203
+ if isinstance(entries, list):
204
+ for entry in entries:
205
+ candidates.extend(str(path) for path in _files_from_info(entry, started_at))
206
+ output: list[Path] = []
207
+ seen: set[str] = set()
208
+ for candidate in candidates:
209
+ path = _safe_download_path(candidate)
210
+ if path and path.is_file() and path.suffix.lower() not in {".part", ".ytdl"} and str(path) not in seen:
211
+ output.append(path)
212
+ seen.add(str(path))
213
+ root = _download_root()
214
+ try:
215
+ for path in root.rglob("*"):
216
+ if path.is_file() and path.suffix.lower() not in {".part", ".ytdl", ".json", ".description", ".vtt", ".srt", ".ass"} and path.stat().st_mtime >= started_at and str(path) not in seen:
217
+ output.append(path)
218
+ seen.add(str(path))
219
+ except OSError:
220
+ pass
221
+ return output
222
+
223
+
224
+ def _operation_id(url: str) -> str:
225
+ digest = hashlib.sha256(f"{url}|{_now()}|{uuid.uuid4().hex}".encode("utf-8")).hexdigest()[:16]
226
+ return f"media_{digest}"
227
+
228
+
229
+ def _notify(record: dict[str, Any]) -> None:
230
+ suffix = "concluído" if record.get("status") == "completed" else "falhou"
231
+ event_type = "media_download_completed" if record.get("status") == "completed" else "media_download_failed"
232
+ title = str(record.get("title") or record.get("display_url") or "Download de mídia")
233
+ message = f"O download de mídia {suffix}."
234
+ if record.get("status") == "failed" and record.get("error"):
235
+ message = f"O download de mídia falhou: {_redact(record['error'])}"
236
+ record_notification(
237
+ event_type,
238
+ title,
239
+ message,
240
+ metadata={"operation_id": record.get("operation_id"), "mode": record.get("mode"), "files": record.get("files", []), "display_url": record.get("display_url")},
241
+ dedupe_key=f"media:{record.get('operation_id')}:{record.get('status')}",
242
+ )
243
+
244
+
245
+ def download_media(
246
+ urls: str | Iterable[str],
247
+ *,
248
+ mode: str = "video",
249
+ quality: str = "Melhor qualidade",
250
+ container: str = "mp4",
251
+ audio_format: str = "mp3",
252
+ allow_playlist: bool = False,
253
+ download_subtitles: bool = False,
254
+ embed_metadata: bool = False,
255
+ progress_callback: ProgressCallback | None = None,
256
+ ) -> list[dict[str, Any]]:
257
+ """Download one or more public URLs and return one persisted record per URL."""
258
+ normalized_urls = normalize_urls(urls)
259
+ results: list[dict[str, Any]] = []
260
+ for url in normalized_urls:
261
+ operation_id = _operation_id(url)
262
+ record: dict[str, Any] = {
263
+ "operation_id": operation_id,
264
+ "url": _display_url(url),
265
+ "display_url": _display_url(url),
266
+ "mode": str(mode or "video").lower(),
267
+ "status": "processing",
268
+ "title": "",
269
+ "files": [],
270
+ "progress": 0.0,
271
+ "created_at": _now(),
272
+ "completed_at": "",
273
+ "error": "",
274
+ }
275
+ _upsert_history(record)
276
+ if progress_callback:
277
+ progress_callback({**record, "status": "processing"})
278
+ started_at = datetime.now().timestamp()
279
+
280
+ def progress_hook(payload: dict[str, Any]) -> None:
281
+ status = str(payload.get("status") or "")
282
+ downloaded = float(payload.get("downloaded_bytes") or 0)
283
+ total = float(payload.get("total_bytes") or payload.get("total_bytes_estimate") or 0)
284
+ progress = min(99.0, (downloaded / total) * 100) if total > 0 else (50.0 if status == "downloading" else 0.0)
285
+ record["progress"] = round(progress, 1)
286
+ if payload.get("filename"):
287
+ record["current_file"] = Path(str(payload["filename"])).name
288
+ if progress_callback:
289
+ progress_callback({**record, "hook_status": status})
290
+
291
+ try:
292
+ if yt_dlp is None:
293
+ raise MediaDownloadError("yt-dlp não está instalado. Instale as dependências do Thunderbolt e tente novamente.")
294
+ options = build_download_options(mode=mode, quality=quality, container=container, audio_format=audio_format, allow_playlist=allow_playlist, download_subtitles=download_subtitles, embed_metadata=embed_metadata, progress_hook=progress_hook)
295
+ downloader = yt_dlp.YoutubeDL(options)
296
+ info = downloader.extract_info(url, download=True)
297
+ close = getattr(downloader, "close", None)
298
+ if callable(close):
299
+ close()
300
+ files = _files_from_info(info, started_at)
301
+ if not files:
302
+ raise MediaDownloadError("O yt-dlp terminou sem produzir um ficheiro local verificável.")
303
+ record.update({"status": "completed", "title": str(info.get("title") or "Download concluído") if isinstance(info, dict) else "Download concluído", "files": [_relative_name(path) for path in files], "progress": 100.0, "completed_at": _now(), "error": ""})
304
+ _upsert_history(record)
305
+ _notify(record)
306
+ except Exception as exc: # the UI receives a persisted failed record per URL
307
+ record.update({"status": "failed", "error": _redact(exc), "completed_at": _now()})
308
+ _upsert_history(record)
309
+ _notify(record)
310
+ results.append(dict(record))
311
+ if progress_callback:
312
+ progress_callback(dict(record))
313
+ return results
@@ -26,6 +26,8 @@ EVENT_CATALOG: tuple[dict[str, str], ...] = (
26
26
  {"code": "cuts_completed", "category": "Edição", "label": "Cortes concluídos", "description": "Quando a geração de cortes terminar com um manifesto completo."},
27
27
  {"code": "metadata_cleaning_completed", "category": "Edição", "label": "Metadados limpos", "description": "Quando uma cópia com metadados limpos for criada."},
28
28
  {"code": "python_edit_completed", "category": "Edição", "label": "Edição Python concluída", "description": "Quando uma operação do Editor Python guardar o artefacto."},
29
+ {"code": "media_download_completed", "category": "Edição", "label": "Download Mídia concluído", "description": "Quando um vídeo ou áudio terminar de ser descarregado com sucesso."},
30
+ {"code": "media_download_failed", "category": "Edição", "label": "Download Mídia falhou", "description": "Quando um download de vídeo ou áudio terminar com erro."},
29
31
  {"code": "automation_completed", "category": "Automação", "label": "Automação concluída", "description": "Quando o worker concluir o lote agendado de um canal."},
30
32
  {"code": "automation_failed", "category": "Automação", "label": "Automação falhou", "description": "Quando uma execução automática terminar com erro."},
31
33
  {"code": "activity_failed", "category": "Sistema", "label": "Actividade falhou", "description": "Quando uma tarefa ou operação persistida terminar em erro."},
@@ -13,6 +13,7 @@ STORAGE = Path(os.getenv("THUNDERBOLT_STORAGE_DIR") or ROOT / "storage")
13
13
  STATE = STORAGE / "state"
14
14
  BLUEPRINTS = STORAGE / "blueprints"
15
15
  TIKTOK_PROMPT_MASTERS = STORAGE / "tiktok" / "prompts_master"
16
+ MEDIA_DOWNLOADS = STORAGE / "downloads"
16
17
  NICHES_DATA = STORAGE / "data" / "niches"
17
18
  SEED_BLUEPRINTS = ROOT / "seed" / "blueprints"
18
19
  SEED_TIKTOK_PROMPT_MASTERS = ROOT / "seed" / "prompt_masters"
@@ -25,6 +26,7 @@ DEFAULTS: dict[str, Any] = {
25
26
  "batches.json": [],
26
27
  "uploads.json": [],
27
28
  "notifications.json": [],
29
+ "media_downloads.json": [],
28
30
  "display_names.json": {"blueprints": {}, "prompt_masters": {}},
29
31
  "niche_apify_runs.json": [],
30
32
  "metadata_edits.json": [],
@@ -283,7 +285,7 @@ def seed_prompt_masters() -> None:
283
285
 
284
286
 
285
287
  def ensure_storage() -> None:
286
- for path in [STATE, BLUEPRINTS / "canais", BLUEPRINTS / "nichos", BLUEPRINTS / "importados", BLUEPRINTS / "brandings", TIKTOK_PROMPT_MASTERS, STORAGE / "brand", STORAGE / "scripts", STORAGE / "thumbnails", STORAGE / "videos", STORAGE / "artifacts", STORAGE / "skills", STORAGE / "metadata_cleaner", STORAGE / "metadata_cleaner" / "outputs", STORAGE / "music", STORAGE / "voice_previews", STORAGE / "python_editor", NICHES_DATA]:
288
+ for path in [STATE, BLUEPRINTS / "canais", BLUEPRINTS / "nichos", BLUEPRINTS / "importados", BLUEPRINTS / "brandings", TIKTOK_PROMPT_MASTERS, MEDIA_DOWNLOADS, STORAGE / "brand", STORAGE / "scripts", STORAGE / "thumbnails", STORAGE / "videos", STORAGE / "artifacts", STORAGE / "skills", STORAGE / "metadata_cleaner", STORAGE / "metadata_cleaner" / "outputs", STORAGE / "music", STORAGE / "voice_previews", STORAGE / "python_editor", NICHES_DATA]:
287
289
  path.mkdir(parents=True, exist_ok=True)
288
290
  seed_blueprints()
289
291
  seed_prompt_masters()
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@danhachuel/thunderbolt",
3
- "version": "0.2.74",
3
+ "version": "0.2.75",
4
4
  "description": "Thunderbolt — interface local para operação de canais faceless e motor MoneyPrinterTurbo",
5
5
  "license": "MIT",
6
6
  "main": "scripts/cli.mjs",
package/requirements.txt CHANGED
@@ -11,5 +11,6 @@ scikit-learn>=1.3,<2
11
11
  mlxtend>=0.23,<1
12
12
  plotly>=5.18,<7
13
13
  seaborn>=0.13,<1
14
+ yt-dlp>=2025.1.15,<2027
14
15
  matplotlib>=3.7,<4
15
16
  kagglehub>=0.3,<1