termux-tts 1.5.0 → 1.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/doc.config.yaml DELETED
@@ -1,595 +0,0 @@
1
- # ==============================================================================
2
- # AMEVA Documentation Configuration: Termux-TTS
3
- # Strict Compliance with AOSF-ENG-STD-2026-V1 & OpenSSF Standards
4
- # Single Source of Truth (SSOT)
5
- # ==============================================================================
6
-
7
- name: "termux-tts"
8
- display_name: "Termux-TTS"
9
- package_name_pypi: "termux-tts"
10
- package_name_npm: "termux-tts"
11
- version: "v1.5.0"
12
- release_name: "100% Native Vulkan Hardware GPU Pipeline & Zero-Silent-Fallback Architecture"
13
- license: "Apache-2.0"
14
- platform: "Android ARM64 / Qualcomm Adreno & ARM Mali Vulkan 1.3 / Linux"
15
- github_repo_url: "https://github.com/uno-km/termux-tts"
16
-
17
- custom_pages:
18
- - slug: "vulkan-engineering-paper"
19
- title_en: "Vulkan C++ Engineering Paper"
20
- title_ko: "Vulkan C++ 네이티브 가속 논문"
21
- file: "lib/tts/vulkan-engineering-paper.html"
22
-
23
- tagline_en: "Production-Grade 4-Tier On-Device Speech Synthesis Framework (Multilingual Neural Orchestrator, Zero-Dependency DSP Formant & C++ Acceleration)"
24
- tagline_ko: "모바일 및 엣지 환경을 위한 프로덕션급 4-Tier 온디바이스 음성 합성 프레임워크 (다국어 신경망 오케스트레이터, 무의존성 DSP 포먼트, C++ 가속 엔진)"
25
-
26
- why_challenge_en: "Constrained mobile edge environments frequently suffer from execution instability, excessive thermal throttling, and unpredictable runtime memory spikes when running conventional deep learning text-to-speech stacks. Heavyweight dependencies like full PyTorch or unoptimized runtime graph compilers exhaust mobile DRAM, while cross-language code-switching historically required multiple disconnected runtimes or heavy cloud APIs."
27
- why_challenge_ko: "제약된 모바일 엣지 환경에서는 과도한 메모리 점유, 발열 스로틀링, 드라이버 파편화로 인해 생성형 음성 합성 모델 구동이 불안정합니다. 기존 무거운 프레임워크는 수 기가바이트의 런타임 의존성으로 모바일 OOM을 유발하며, 다국어 교차 발화(한/영/일 등)는 단일 기기 로컬에서 매끄럽게 처리하기 어려웠습니다."
28
-
29
- description_en: "Termux-TTS delivers a resilient 4-Tier on-device text-to-speech architecture designed for deterministic latency, zero-config ergonomics, and multi-language acoustic fidelity. It bridges lightweight parametric DSP synthesis (<50ms compute, 0MB model download) with high-fidelity resident C-API multilingual neural models (9 languages including Korean, English, Japanese, Hindi, and Russian at 22.05kHz), alongside Android system native speech service routing. With automated 1-click provisioning, Unicode script classification, and resident in-memory model caching, Termux-TTS achieves high acoustic fidelity with sub-0.18x real-time factor."
30
- description_ko: "Termux-TTS는 결정론적 지연 시간, Zero-Config 사용성, 다국어 고품질 음향을 실현하는 프로덕션급 4-Tier 온디바이스 음성 합성 프레임워크입니다. 0MB 무의존성 파라메트릭 DSP 보코더부터 9개 공식 언어(한국어, 영어, 일본어, 힌디어, 러시아어 등)를 지원하는 레지던트 C-API 다국어 신경망 오케스트레이터, 그리고 안드로이드 네이티브 음성 브릿지까지 통합 제공합니다."
31
-
32
- quick_install_cmd: |
33
- pip install termux-tts
34
- # or: npm install termux-tts
35
- # 1-Click Provision Default Models (Korean KSS + English Lessac):
36
- termux-tts install
37
- # On-demand provision specific languages (e.g. Hindi, Japanese, Russian, or all):
38
- termux-tts install --models hi
39
- # Instant speech synthesis:
40
- termux-tts "Hello 방가방가 나는 parrot 이라고 해." -o out.wav --play
41
-
42
- features:
43
- - title_en: "Multilingual Neural Orchestrator"
44
- desc_en: "Dynamic cross-language code-switching and single-language synthesis supporting 9 official languages (Korean, English, Japanese, Chinese, Hindi, Russian, Spanish, French, German) with 50ms context-aware pause padding."
45
- title_ko: "다국어 신경망 오케스트레이터"
46
- desc_ko: "한국어, 영어, 일본어, 중국어, 힌디어, 러시아어, 스페인어, 프랑스어, 독일어 9대 언어의 다국어 교차 발화 및 단독 발화를 50ms 문맥 묵음 패딩과 함께 실시간 합성합니다."
47
- - title_en: "Resident C-API In-Memory Acceleration"
48
- desc_en: "Native C-API residency with SherpaResidentManager eliminates subprocess startup latency and achieves sub-0.18x real-time factor with ARM NEON SIMD acceleration."
49
- title_ko: "레지던트 C-API 인메모리 가속"
50
- desc_ko: "SherpaResidentManager를 통한 C-API 모델 인메모리 상주로 프로세스 기동 지연을 전면 배제하고 ARM NEON SIMD 가속으로 0.18x 이하의 RTF를 달성합니다."
51
- - title_en: "Zero-Config CLI Ergonomics"
52
- desc_en: "Top-level speech synthesis by default (termux-tts 'Hello' --play) without requiring subcommands or explicit engine flags. Explicit 'speak' command routes to Android native system voice."
53
- title_ko: "Zero-Config CLI 사용성"
54
- desc_ko: "서브 커맨드나 엔진 플래그 없이 문장을 직접 입력(termux-tts '문장' --play)하면 즉시 신경망 합성이 수행되며, 안드로이드 시스템 음성은 'speak'로 명시 분기합니다."
55
- - title_en: "On-Demand Self-Healing Provisioner"
56
- desc_en: "1-Click automated provisioning (termux-tts install --models default/hi/ja/ru/zh/all) with actionable English guidance and absolute paths for uninstalled models."
57
- title_ko: "온디맨드 자가 치유 프로비저너"
58
- desc_ko: "최초 설치 시 기본 모델(한/영)만 경량 배포하고, 미설치 언어 호출 시 절대 경로와 함께 즉시 복구 가능한 영문 안내문 및 온디맨드 설치 명령을 제공합니다."
59
- - title_en: "4-Tier Resilient Architecture"
60
- desc_en: "Tier 1: Zero-Dependency Parametric DSP Formant (0MB footprint). Tier 2: Android Native System Voice Bridge. Tier 3: Subprocess-Isolated Sherpa C++ CPU Engine. Tier 4: Pure Vulkan GPU Hardware Neural Acceleration."
61
- title_ko: "4-Tier 복원형 합성 아키텍처"
62
- desc_ko: "Tier 1 무의존성 DSP 포먼트, Tier 2 안드로이드 네이티브 음성 브릿지, Tier 3 격리형 Sherpa CPU 엔진, Tier 4 순수 Vulkan GPU 하드웨어 신경망 가속."
63
- - title_en: "1-Click Automated Provisioning"
64
- desc_en: "Command-line provisioning (termux-tts install --tier high/medium) automatically resolves precompiled ARM64 native binaries and HuggingFace weights with self-test verification."
65
- title_ko: "1-클릭 자동 프로비저닝"
66
- desc_ko: "명령어 1회로 ARM64 네이티브 바이너리와 허깅페이스 가중치를 자동 감지 및 설치하고 자체 진단 테스트를 완결합니다."
67
- - title_en: "Studio-Grade FP16 Neural Fidelity"
68
- desc_en: "Native support for high-resolution 22.05kHz 16-bit PCM studio models (vits-piper-en_US-lessac-high-fp16), delivering natural prosody on mobile hardware."
69
- title_ko: "스튜디오급 FP16 고음질 신경망"
70
- desc_ko: "22.05kHz 16비트 PCM 고해상도 레퍼런스 모델을 지원하여 극도로 자연스러운 억양과 음향 충실도를 모바일 단말기에서 제공합니다."
71
- - title_en: "Pure Vulkan GPU Acceleration (Zero CPU Fallback)"
72
- desc_en: "Direct GPU tensor compute via Vulkan 1.3 pipeline caching, fully resolving Qualcomm Adreno 830 and ARM Mali-G68 shader driver anomalies without silent CPU fallback."
73
- title_ko: "순수 Vulkan GPU 가속 (Zero CPU 침묵 폴백)"
74
- desc_ko: "침묵 폴백 없이 Qualcomm Adreno 830 및 ARM Mali-G68의 셰이더 드라이버 결함을 완전 해소하고 순수 GPU 텐서 연산을 수행합니다."
75
- - title_en: "Subprocess IPC Memory Isolation"
76
- desc_en: "Architectural separation of audio synthesis and OpenSL ES playback threads via subprocess IPC prevents Python GIL contention and memory corruption."
77
- title_ko: "프로세스 격리형 메모리 보호"
78
- desc_ko: "음성 합성 파이프라인과 오디오 재생 스레드를 서브프로세스 IPC로 분리하여 GIL 병목 및 메모리 오염을 원천 차단합니다."
79
- - title_en: "Unified Cross-Language Toolchain"
80
- desc_en: "Identical functional parity and strict typing across Global CLI, Python 3.10+ SDK, and Node.js / TypeScript runtime bindings."
81
- title_ko: "통합 크로스 랭귀지 툴체인"
82
- desc_ko: "Global CLI, Python 3.10+ SDK, Node.js / TypeScript 런타임 간 완전한 기능 동등성과 타입 안전성을 보장합니다."
83
-
84
- matrix_table:
85
- - category: "Tier 1: DSP Formant"
86
- operations: "Rosenberg Glottal Pulse, 5-Band Biquad Filter, Sino-Korean Numeral Normalizer"
87
- status: "Production (<50ms)"
88
- - category: "Tier 2: Native Bridge"
89
- operations: "Termux-API / Android TTS Service IPC (Samsung/Google Engine)"
90
- status: "Production (Immediate)"
91
- - category: "Tier 3: CPU Neural"
92
- operations: "Sherpa C++ Subprocess-Isolated VITS Engine (ARM64 NEON)"
93
- status: "Production (RTF 0.4~1.2x)"
94
- - category: "Tier 4: Vulkan GPU"
95
- operations: "C++ Vulkan Pipeline Caching, Mali Quirk Handling, Adreno Subgroup 64"
96
- status: "Production (RTF 0.26x)"
97
-
98
- code_example_py: |
99
- import termux_tts as tts
100
-
101
- # 1. Zero-Config Multilingual Neural Synthesis (Korean + English Code-Switching)
102
- with tts.load() as engine:
103
- result = engine.synthesize(
104
- "Hello 방가방가 키키키키 나는 parrot 이라고 해. Natural cross-language neural voice.",
105
- output="multilingual.wav"
106
- )
107
- print(f"Synthesized {result.duration_sec:.2f}s in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
108
-
109
- # 2. Pure Vulkan GPU Neural Synthesis (Studio Tier)
110
- with tts.load(engine="vulkan", model_tier="high") as engine:
111
- result = engine.synthesize(
112
- "The neural speech synthesis engine is operating with pure Vulkan hardware acceleration.",
113
- output="studio_vulkan.wav"
114
- )
115
- print(f"Synthesized {result.duration_sec:.2f}s in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
116
-
117
- # 3. Instant Zero-Dependency DSP Formant Synthesis
118
- with tts.load(engine="dsp", preset="balanced") as engine:
119
- result = engine.synthesize("Instant speech generation with zero external weights.", output="dsp.wav")
120
- print(f"DSP Synthesis Latency: {result.elapsed_ms:.1f}ms")
121
-
122
- # 4. Direct Android Hardware Speaker Playback
123
- with tts.load(engine="native", language="en") as engine:
124
- engine.speak("Direct hardware speaker output via Android native service.")
125
-
126
- code_example_js: |
127
- const tts = require('termux-tts');
128
-
129
- async function main() {
130
- // 1. Initialize Vulkan GPU Neural Engine
131
- const engine = tts.load({ engine: 'vulkan', tier: 'high' });
132
- const res = await engine.synthesize(
133
- "High-performance speech synthesis on mobile hardware.",
134
- { output: "vulkan_output.wav" }
135
- );
136
- console.log(`Generated: ${res.durationSec}s in ${res.elapsedMs}ms (RTF: ${res.rtf}x)`);
137
-
138
- // 2. Hardware Diagnostics
139
- const diag = await tts.doctor();
140
- console.log(`Vulkan GPU Device: ${diag.device_name} (API: ${diag.api_version})`);
141
- }
142
- main();
143
-
144
- benchmarks:
145
- headers: ["Hardware Target", "SoC / Physical GPU Architecture", "Vulkan Driver ABI", "Audio Duration", "Synthesis Latency", "Real-Time Factor (RTF)", "Execution Status"]
146
- rows:
147
- - ["Galaxy S21", "Exynos 2100 / ARM Mali-G78 (0x9800000)", "Vulkan 1.1 (/system/lib64)", "4.25 s", "9.37 s", "2.2055x", "Validated (100% Native GPU)"]
148
- - ["Galaxy S25", "Snapdragon 8 Elite / Adreno 830 (0x80320040)", "Vulkan 1.3 (/system/lib64)", "4.25 s", "16.21 s", "3.8145x", "Validated (100% Native GPU)"]
149
- - ["Galaxy S22", "Snapdragon 8 Gen 1 / Adreno 730 (0x80267062)", "Vulkan 1.1 (/system/lib64)", "4.27 s", "38.09 s", "8.9159x", "Validated (100% Native GPU)"]
150
- - ["Galaxy A35", "Exynos 1380 / ARM Mali-G68 MP5 (0x9801000)", "Vulkan 1.1 (/system/lib64)", "4.27 s", "51.91 s", "12.1505x", "Validated (100% Native GPU)"]
151
- - ["Heterogeneous CPU", "Cortex-A78 / A55 CPU Core", "Sherpa NEON SIMD C++", "4.25 s", "1.85 s", "0.4350x", "Validated (CPU Reference)"]
152
-
153
- api_reference:
154
- - symbol: "termux_tts.load(engine='auto'|'vulkan'|'sherpa'|'dsp'|'native', model_tier='high'|'medium', preset='balanced')"
155
- description: "Factory initializer returning the configured synthesis engine context with resource finalization."
156
- - symbol: "engine.synthesize(text: str, output: str = None, speed: float = 1.0, preset: str = None) -> SynthesisResult"
157
- description: "Executes phonetic tokenization, neural/DSP acoustic modeling, and writes 16-bit PCM WAV audio."
158
- - symbol: "engine.speak(text: str, stream: str = 'MUSIC') -> NativeResult"
159
- description: "Routes speech synthesis directly to the physical device speakers via Android system native services."
160
- - symbol: "termux_tts.run_installation(tier='high'|'medium', force=False) -> bool"
161
- description: "Automated provisioning toolchain downloading native C++ binaries from GitHub Releases and HuggingFace weights."
162
- - symbol: "termux_tts.doctor() -> DiagnosticReport"
163
- description: "Runs rigorous 12-stage hardware diagnostics verifying Vulkan instance, GPU compute queues, and memory buffers."
164
-
165
- installation_body: |
166
- <h3>Standard Package Installation</h3>
167
- <p>Install the core Python or Node.js package directly from standard package repositories:</p>
168
- <pre><code># Python (PyPI)
169
- pip install termux-tts
170
-
171
- # Node.js / TypeScript (NPM)
172
- npm install termux-tts</code></pre>
173
-
174
- <h3>Prerequisites on Android Termux</h3>
175
- <p>Ensure that the required system packages are installed in your Termux environment:</p>
176
- <pre><code>pkg update && pkg install -y termux-api pulseaudio sox clang</code></pre>
177
-
178
- <h3>1-Click Automated Engine Provisioning</h3>
179
- <p>To enable Tier 4 Vulkan GPU neural speech synthesis without manual compilation or model downloading, run the built-in installer:</p>
180
- <pre><code># Install High-Resolution Studio Tier (vits-piper-en_US-lessac-high-fp16, 22.05kHz)
181
- termux-tts install --tier high
182
-
183
- # Or install Medium Compact Tier (vits-piper-en_US-lessac-medium)
184
- termux-tts install --tier medium</code></pre>
185
- <p>The provisioning script downloads the precompiled native ARM64 C++ binary (<code>sherpa-ncnn-offline-tts-vulkan</code>) from the official GitHub Release and verifies execution via an immediate self-test.</p>
186
-
187
- quickstart_body: |
188
- <h3>Command-Line Interface (CLI) Recipes</h3>
189
- <pre><code># 1. High-Resolution Studio Neural Synthesis with Speaker Playback
190
- termux-tts synth -e vulkan --tier high -t "Operating at full hardware capacity." -o speech.wav --play
191
-
192
- # 2. Instant Zero-Dependency DSP Formant Synthesis
193
- termux-tts synth -e dsp -p ultra -t "Zero dependency parametric speech synthesis." -o dsp.wav
194
-
195
- # 3. Direct Android Hardware Speaker Broadcast
196
- termux-tts speak -t "System notification broadcast." -l en
197
-
198
- # 4. Hardware and Environment Diagnostics
199
- termux-tts doctor</code></pre>
200
-
201
- <h3>Python SDK Canon</h3>
202
- <pre><code>import termux_tts as tts
203
-
204
- # Initialize pure Vulkan GPU neural synthesis engine
205
- with tts.load(engine="vulkan", model_tier="high") as engine:
206
- res = engine.synthesize("Validating deterministic tensor execution.", output="out.wav")
207
- print(f"Elapsed: {res.elapsed_ms:.1f}ms | RTF: {res.rtf:.4f}x")</code></pre>
208
-
209
- api_body: |
210
- <h3>Factory Gateway</h3>
211
- <table class="data-table">
212
- <thead>
213
- <tr><th>Symbol</th><th>Signature</th><th>Description</th></tr>
214
- </thead>
215
- <tbody>
216
- <tr>
217
- <td><code>termux_tts.load</code></td>
218
- <td><code>(engine: str = 'auto', model_tier: str = 'high', preset: str = 'balanced', language: str = 'en')</code></td>
219
- <td>Initializes and returns the unified TTS engine instance. Supports context manager pattern.</td>
220
- </tr>
221
- <tr>
222
- <td><code>termux_tts.run_installation</code></td>
223
- <td><code>(tier: str = 'high', force: bool = False) -> bool</code></td>
224
- <td>1-click provisioning tool downloading native ARM64 C++ binaries and model weights.</td>
225
- </tr>
226
- <tr>
227
- <td><code>termux_tts.doctor</code></td>
228
- <td><code>() -> DiagnosticReport</code></td>
229
- <td>Executes comprehensive 12-stage Vulkan and system diagnostics.</td>
230
- </tr>
231
- </tbody>
232
- </table>
233
-
234
- <h3>Engine Methods</h3>
235
- <table class="data-table">
236
- <thead>
237
- <tr><th>Method</th><th>Parameters</th><th>Return Type</th><th>Description</th></tr>
238
- </thead>
239
- <tbody>
240
- <tr>
241
- <td><code>synthesize</code></td>
242
- <td><code>(text: str, output: str = None, speed: float = 1.0, preset: str = None)</code></td>
243
- <td><code>SynthesisResult</code></td>
244
- <td>Synthesizes text into 16-bit linear PCM WAV audio.</td>
245
- </tr>
246
- <tr>
247
- <td><code>speak</code></td>
248
- <td><code>(text: str, stream: str = 'MUSIC')</code></td>
249
- <td><code>NativeResult</code></td>
250
- <td>Transmits speech to the device speaker via Android native TTS bridge.</td>
251
- </tr>
252
- </tbody>
253
- </table>
254
-
255
- benchmarks_body: |
256
- <h3>Empirical Dual-Device Hardware Benchmark Suite</h3>
257
- <p>The following performance benchmarks were gathered directly on physical test devices running Android 16 under Termux ARM64 environment. Every test was conducted over repeated synthesis iterations with zero silent fallback.</p>
258
-
259
- <table class="data-table">
260
- <thead>
261
- <tr>
262
- <th>Device &amp; Hardware Profile</th>
263
- <th>Synthesis Engine</th>
264
- <th>Model Architecture</th>
265
- <th>Audio Duration</th>
266
- <th>Compute Latency</th>
267
- <th>Real-Time Factor</th>
268
- <th>Hardware Utilization</th>
269
- </tr>
270
- </thead>
271
- <tbody>
272
- <tr>
273
- <td><strong>Samsung Galaxy S25</strong><br>Snapdragon 8 Elite / Adreno 830</td>
274
- <td>Vulkan GPU Neural</td>
275
- <td>lessac-high-fp16 (22.05kHz)</td>
276
- <td>6.70 s</td>
277
- <td><strong>6.65 s</strong></td>
278
- <td><strong>0.993x</strong></td>
279
- <td>GPU Subgroup 64 / Pure Tensor</td>
280
- </tr>
281
- <tr>
282
- <td><strong>Samsung Galaxy S25</strong><br>Snapdragon 8 Elite / Adreno 830</td>
283
- <td>Vulkan GPU Neural</td>
284
- <td>lessac-medium (22.05kHz)</td>
285
- <td>4.59 s</td>
286
- <td><strong>1.21 s</strong></td>
287
- <td><strong>0.264x</strong></td>
288
- <td>GPU Compute / 3.79x Faster Than RT</td>
289
- </tr>
290
- <tr>
291
- <td><strong>Samsung Galaxy A35</strong><br>Exynos 1380 / ARM Mali-G68 MP5</td>
292
- <td>Vulkan GPU Neural</td>
293
- <td>lessac-medium (22.05kHz)</td>
294
- <td>4.52 s</td>
295
- <td><strong>5.18 s</strong></td>
296
- <td><strong>1.146x</strong></td>
297
- <td>Mali Subgroup 16 / Medium Kernel</td>
298
- </tr>
299
- <tr>
300
- <td><strong>Samsung Galaxy A35</strong><br>Exynos 1380 / ARM Mali-G68 MP5</td>
301
- <td>Vulkan GPU Neural</td>
302
- <td>lessac-high-fp16 (22.05kHz)</td>
303
- <td>6.73 s</td>
304
- <td><strong>34.33 s</strong></td>
305
- <td><strong>5.098x</strong></td>
306
- <td>Full VRAM Resident / High Fidelity</td>
307
- </tr>
308
- <tr>
309
- <td><strong>Heterogeneous ARM64</strong><br>All Core Profiles</td>
310
- <td>Parametric DSP Formant</td>
311
- <td>5-Band Biquad Formant</td>
312
- <td>4.15 s</td>
313
- <td><strong>0.054 s</strong></td>
314
- <td><strong>0.0130x</strong></td>
315
- <td>0MB External RAM / Instant CPU</td>
316
- </tr>
317
- </tbody>
318
- </table>
319
-
320
- advanced_parameters_body: |
321
- <h3>Architecture &amp; Hardware Driver Postmortem</h3>
322
- <p>Operating deep learning inference models on mobile Vulkan runtimes involves navigating severe GPU driver irregularities. Below is the technical breakdown of bugs identified and permanently resolved in Termux-TTS v1.3.0:</p>
323
-
324
- <h4>1. ARM Mali-G68 Valhall Subgroup Truncation Resolution</h4>
325
- <p>On ARM Mali-G68 GPUs (subgroup size 16), unaligned matrix multiplication workgroups caused an integer division truncation: <code>loadstride_b = gl_WorkGroupSize.x * LOAD_VEC_B / BK = 16 * 1 / 32 = 0</code>. This resulted in an infinite loop (<code>for (uint l = 0; l &lt; BN; l += 0)</code>) leading to GPU command timeout and device loss (<code>VK_ERROR_DEVICE_LOST</code>). Termux-TTS enforces medium tile alignment and subgroup-aware shader dispatch, eliminating the loop and ensuring full 100% layer execution.</p>
326
-
327
- <h4>2. Qualcomm Adreno 830 SPIR-V Pipeline Caching</h4>
328
- <p>On Snapdragon 8 Elite Adreno 830 hardware, repeated runtime pipeline generation triggered JIT compiler crashes (<code>VK_ERROR_UNKNOWN -13</code>). Termux-TTS isolates model pipelines with immutable descriptor pooling and shader pipeline caching, stabilizing execution across long synthesis streams.</p>
329
-
330
- <h4>3. Subprocess IPC Architecture for Audio Playback</h4>
331
- <p>In-process audio playback via native libraries frequently triggered memory corruption and GIL deadlocks on Android Bionic. Termux-TTS orchestrates playback using isolated subprocess worker pools, safeguarding the primary application runtime.</p>
332
-
333
- changelog:
334
- - version: "v1.4.4"
335
- date: "2026-09-14"
336
- title: "Multilingual Neural Orchestrator & Zero-Config Ergonomics"
337
- type: "Production Release (Latest)"
338
- changes:
339
- - "Implemented MultilingualNeuralEngine supporting dynamic cross-language code-switching across 9 languages (ko, en, ja, zh, hi, ru, es, fr, de)."
340
- - "Built SherpaResidentManager for native in-memory C-API residency with sub-0.18x RTF and zero subprocess startup lag."
341
- - "Promoted speech synthesis to default top-level subcommand, accepting positional text without '-t' or '-e multilingual'."
342
- - "Integrated universal Unicode script classifier covering Hangul, Latin, Devanagari, Cyrillic, CJK, and Arabic scripts."
343
- - "Enhanced 1-click installer with lightweight default models (ko+en) and on-demand provisioning (--models <lang> / all)."
344
- - "Hardened defensive path handling with Path.expanduser() across audio exports and CLI arguments."
345
-
346
- - version: "v1.3.0"
347
- date: "2026-09-05"
348
- title: "Production 4-Tier Synthesizer & Studio Vulkan GPU Acceleration"
349
- type: "Stable Release"
350
- changes:
351
- - "Implemented Tier 4 Pure Vulkan GPU Neural Engine (sherpa-ncnn-offline-tts-vulkan) with zero CPU fallback."
352
- - "Integrated studio-grade reference model (vits-piper-en_US-lessac-high-fp16, 22.05kHz 16-bit PCM)."
353
- - "Implemented 1-click automated provisioning toolchain (termux-tts install --tier high/medium) downloading ARM64 native binaries from GitHub Release."
354
- - "Resolved ARM Mali-G68 subgroup size 16 shader integer truncation and Adreno 830 pipeline cache crashes."
355
- - "Achieved empirical RTF 0.264x on Galaxy S25 Adreno 830 and RTF 1.146x on Galaxy A35 Mali-G68."
356
- - "Unified cross-platform CLI, Python SDK, and Node.js / TypeScript runtime bindings."
357
-
358
- - version: "v1.1.5"
359
- date: "2026-09-04"
360
- title: "Subprocess Isolation & Conversational Tag Expansion"
361
- type: "Archive Release"
362
- changes:
363
- - "Implemented subprocess IPC audio playback isolation."
364
- - "Added expressive conversational token handling."
365
- - "Stabilized Termux-API native bridge integration."
366
-
367
- readme_content: |
368
- # Termux-TTS
369
-
370
- [![PyPI](https://img.shields.io/pypi/v/termux-tts.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-tts/)
371
- [![Python](https://img.shields.io/pypi/pyversions/termux-tts.svg?style=flat-square)](https://pypi.org/project/termux-tts/)
372
- [![npm](https://img.shields.io/npm/v/termux-tts.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-tts)
373
- [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-tts)
374
-
375
- > Production-Grade 4-Tier On-Device Speech Synthesis Framework (Multilingual Neural Orchestrator, Zero-Dependency DSP Formant & C++ Acceleration)
376
-
377
- ---
378
-
379
- ## Architecture & Overview
380
-
381
- Termux-TTS is an enterprise-grade, on-device text-to-speech framework optimized for mobile edge hardware and Android Termux environments. Built to eliminate heavy dependency stacks and fragile driver behaviors, it features a resilient 4-Tier architecture:
382
-
383
- - **Multilingual Neural Orchestrator**: Dynamic cross-language code-switching mesh covering 9 official languages (Korean, English, Japanese, Chinese, Hindi, Russian, Spanish, French, German) with Unicode script tokenization and 50ms context-aware silence padding.
384
- - **Resident C-API In-Memory Engine**: Direct C-API memory residency with `SherpaResidentManager` achieving sub-0.18x real-time factor with ARM NEON SIMD acceleration and zero subprocess lag.
385
- - **Tier 1: Zero-Dependency Parametric DSP Formant Vocoder**: 0MB disk footprint, Rosenberg glottal pulse formulation, and 5-band biquad formant filters providing deterministic speech synthesis in under 50 milliseconds (RTF 0.013x).
386
- - **Tier 2: Android System Native Voice Bridge**: Direct IPC integration to physical Samsung and Google speech engines via the Termux-API service layer.
387
- - **Tier 3: Subprocess-Isolated Sherpa C++ CPU Engine**: Subprocess-isolated VITS acoustic modeling on ARM64 NEON with memory leak protection.
388
- - **Tier 4: Pure Vulkan GPU Hardware Neural Engine**: High-performance GPU tensor synthesis via precompiled native C++ binaries (`sherpa-ncnn-offline-tts-vulkan`) running high-resolution studio models (`vits-piper-en_US-lessac-high-fp16`) with zero silent CPU fallback.
389
-
390
- ---
391
-
392
- ## Empirical Hardware Benchmarks (Physical Devices)
393
-
394
- Measurements gathered on physical Android 16 hardware running Termux ARM64:
395
-
396
- | Target Device | Hardware Architecture | Synthesis Engine | Model Profile | Audio Length | Synthesis Time | Real-Time Factor (RTF) | Status |
397
- | :--- | :--- | :--- | :--- | :--- | :---: | :---: | :---: |
398
- | **Galaxy A53** | Exynos 1280 / ARM64 NEON | Multilingual Neural C-API | `KSS + Lessac` (Bilingual) | 10.50 s | **1.82 s** | **0.173x** | Validated (5.78x RT) |
399
- | **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `lessac-high-fp16` | 6.70 s | **6.65 s** | **0.993x** | Validated |
400
- | **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `lessac-medium` | 4.59 s | **1.21 s** | **0.264x** | Validated (3.79x RT) |
401
- | **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU Neural | `lessac-medium` | 4.52 s | **5.18 s** | **1.146x** | Validated |
402
- | **ARM64 CPU** | Cortex-A78 / A55 | Parametric DSP | 5-Band Biquad | 4.15 s | **0.054 s** | **0.0130x** | Validated (76x RT) |
403
-
404
- ---
405
-
406
- ## Installation & Automated Provisioning
407
-
408
- ### 1. Package Installation
409
- ```bash
410
- # Python SDK & CLI
411
- pip install termux-tts
412
-
413
- # Node.js / TypeScript
414
- npm install termux-tts
415
- ```
416
-
417
- ### 2. Automated Model Provisioning
418
- Automate downloading and linking precompiled models to the official immutable path (`/data/data/com.termux/files/home/models/tts/`):
419
- ```bash
420
- # Default lightweight installation (Korean KSS + English Lessac)
421
- termux-tts install
422
-
423
- # On-demand provision specific language models:
424
- termux-tts install --models hi # Hindi (Piper Swara)
425
- termux-tts install --models ja # Japanese (Piper Hina)
426
- termux-tts install --models ru # Russian (Piper Dmitri)
427
- termux-tts install --models zh # Chinese (AISHELL3)
428
- termux-tts install --models all # All 9 official languages
429
-
430
- # Provision Studio Vulkan GPU engine:
431
- termux-tts install --tier high
432
- ```
433
-
434
- ---
435
-
436
- ## Quickstart
437
-
438
- ### Global Command-Line Interface (Zero-Config)
439
- ```bash
440
- # 1. Zero-config synthesis with physical speaker playback
441
- termux-tts "Hello 방가방가 나는 parrot 이라고 해." -o ~/out.wav --play
442
-
443
- # 2. Flag-based syntax
444
- termux-tts -t "Multilingual neural speech synthesis on device." --play
445
-
446
- # 3. Force language pinning
447
- termux-tts -l en -t "Pure English text output." --play
448
- termux-tts -l ko -t "한국어 단독 신경망 음성 합성." --play
449
-
450
- # 4. Instant DSP Formant synthesis (0MB footprint)
451
- termux-tts synth -e dsp -t "Zero dependency DSP synthesis." -o dsp.wav
452
-
453
- # 5. Direct Android system native speaker broadcast
454
- termux-tts speak -t "Hardware speaker broadcast via Android service."
455
-
456
- # 6. Hardware diagnostics
457
- termux-tts doctor
458
- ```
459
-
460
- ### Python SDK
461
- ```python
462
- import termux_tts as tts
463
-
464
- # 1. Zero-Config Multilingual Neural Synthesis (Korean + English Code-Switching)
465
- with tts.load() as engine:
466
- result = engine.synthesize(
467
- "Hello 방가방가 키키키키 나는 parrot 이라고 해.",
468
- output="multilingual.wav"
469
- )
470
- print(f"Generated {result.duration_sec:.2f}s in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
471
-
472
- # 2. Pure Vulkan GPU Neural Synthesis (Studio Tier)
473
- with tts.load(engine="vulkan", model_tier="high") as engine:
474
- result = engine.synthesize("Pure Vulkan neural execution on mobile.", output="speech.wav")
475
- print(f"Synthesized in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
476
-
477
- # 3. Zero-Dependency DSP Formant Synthesis
478
- with tts.load(engine="dsp", preset="balanced") as engine:
479
- result = engine.synthesize("Instant speech without model downloads.", output="dsp.wav")
480
- print(f"DSP Latency: {result.elapsed_ms:.1f}ms")
481
- ```
482
-
483
- ### Node.js / TypeScript
484
- ```typescript
485
- import * as tts from 'termux-tts';
486
-
487
- async function main() {
488
- const engine = tts.load({ engine: 'vulkan', tier: 'high' });
489
- const res = await engine.synthesize("High performance speech synthesis.", { output: "speech.wav" });
490
- console.log(`Synthesized in ${res.elapsedMs}ms`);
491
- }
492
- main();
493
- ```
494
-
495
- ---
496
-
497
- ## Official Documentation & Benchmarks
498
- - [Official Architecture & API Reference](https://uno-km.vercel.app/lib/tts/)
499
- - [Ecosystem Metrics & Registry Stats](https://uno-km.vercel.app/foundation/metrics)
500
- - [AMEVA Open-Source Foundation Portal](https://uno-km.vercel.app/foundation/index.html)
501
-
502
- ---
503
-
504
- ## License
505
- Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)).
506
-
507
- readme_pypi_content: |
508
- # Termux-TTS (Python)
509
-
510
- [![PyPI](https://img.shields.io/pypi/v/termux-tts.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-tts/)
511
- [![Python](https://img.shields.io/pypi/pyversions/termux-tts.svg?style=flat-square)](https://pypi.org/project/termux-tts/)
512
- [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-tts)
513
-
514
- > Production-Grade 4-Tier On-Device Speech Synthesis Framework for Mobile & Edge (Zero-Dependency DSP Formant, C++ Vulkan GPU Neural Engine & Android Native Voice Bridge)
515
-
516
- ## Installation
517
-
518
- ```bash
519
- pip install termux-tts
520
- ```
521
-
522
- ### 1-Click Automated Engine Provisioning
523
- ```bash
524
- termux-tts install --tier high
525
- ```
526
-
527
- ## Quickstart
528
-
529
- ```python
530
- import termux_tts as tts
531
-
532
- # 1. Studio Vulkan GPU Neural Engine
533
- with tts.load(engine="vulkan", model_tier="high") as engine:
534
- result = engine.synthesize("Neural speech synthesis on mobile GPU.", output="speech.wav")
535
- print(f"Elapsed: {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
536
-
537
- # 2. Zero-Dependency DSP Formant Mode
538
- with tts.load(engine="dsp", preset="balanced") as engine:
539
- result = engine.synthesize("Instant speech generation.", output="dsp.wav")
540
- print(f"DSP Latency: {result.elapsed_ms:.1f}ms")
541
-
542
- # 3. Direct Android Native Speaker Output
543
- with tts.load(engine="native", language="en") as engine:
544
- engine.speak("Direct hardware speaker output.")
545
- ```
546
-
547
- ## Benchmarks (Physical Devices)
548
-
549
- | Target Device | Hardware Architecture | Synthesis Engine | Real-Time Factor (RTF) | Status |
550
- | :--- | :--- | :--- | :---: | :---: |
551
- | **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU (`lessac-high-fp16`) | **0.993x** | Validated |
552
- | **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU (`lessac-medium`) | **0.264x** | Validated |
553
- | **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU (`lessac-medium`) | **1.146x** | Validated |
554
- | **ARM64 CPU** | All Core Profiles | Parametric DSP Formant | **0.0130x** | Validated |
555
-
556
- ## Documentation
557
- - [Official Documentation & API Reference](https://uno-km.vercel.app/lib/tts/)
558
- - [GitHub Repository](https://github.com/uno-km/termux-tts)
559
-
560
- ## License
561
- Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).
562
-
563
- readme_npm_content: |
564
- # Termux-TTS (Node.js & TypeScript)
565
-
566
- [![npm](https://img.shields.io/npm/v/termux-tts.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-tts)
567
- [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-tts)
568
-
569
- > Production-Grade 4-Tier On-Device Speech Synthesis Framework for Mobile & Edge (Zero-Dependency DSP Formant, C++ Vulkan GPU Neural Engine & Android Native Voice Bridge)
570
-
571
- ## Installation
572
-
573
- ```bash
574
- npm install termux-tts
575
- ```
576
-
577
- ## Quickstart
578
-
579
- ```typescript
580
- import * as tts from 'termux-tts';
581
-
582
- async function main() {
583
- const engine = tts.load({ engine: 'vulkan', tier: 'high' });
584
- const res = await engine.synthesize("High performance speech synthesis on mobile.", { output: "output.wav" });
585
- console.log(`Synthesized: ${res.durationSec}s in ${res.elapsedMs}ms (RTF: ${res.rtf}x)`);
586
- }
587
- main();
588
- ```
589
-
590
- ## Documentation
591
- - [Official Documentation & API Reference](https://uno-km.vercel.app/lib/tts/)
592
- - [GitHub Repository](https://github.com/uno-km/termux-tts)
593
-
594
- ## License
595
- Apache-2.0 License. Copyright (c) 2026 Eunho Kim (@uno-km).
@@ -1,12 +0,0 @@
1
- # 모바일 지연시간 벤치마크 (Mobile Benchmarks)
2
-
3
- ## 1. 실기기 벤치마크 결과
4
-
5
- | 디바이스 | AP 프로세서 | 모델 | RTF (Real-Time Factor) | 5초 문장 합성 소요시간 | 메모리 점유율 |
6
- | :--- | :--- | :--- | :--- | :--- | :--- |
7
- | **Galaxy S25 (SM-S931N)** | Snapdragon 8 Elite | VITS ONNX | **0.035x** | **0.18초** | ~68MB |
8
- | **Galaxy A35 (SM-A356N)** | Exynos 1380 | VITS ONNX | **0.118x** | **0.59초** | ~72MB |
9
- | **Galaxy S20 (SM-G981N)** | Snapdragon 865 | VITS ONNX | **0.142x** | **0.71초** | ~75MB |
10
-
11
- > **RTF 공식**: $\text{RTF} = \frac{\text{합성에 소요된 연산 시간 (초)}}{\text{생성된 오디오 길이 (초)}}$
12
- > RTF가 1.0 미만이면 실시간 발화 속도보다 빠르게 합성됨을 의미합니다.