termux-tts 1.4.2 → 1.4.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +21 -0
- package/README.md +81 -381
- package/README.pypi.md +26 -410
- package/doc.config.yaml +101 -37
- package/package.json +1 -1
- package/pyproject.toml +1 -1
- package/termux_tts/__init__.py +11 -1
- package/termux_tts/audio.py +6 -3
- package/termux_tts/cli.py +27 -10
- package/termux_tts/engine.py +73 -7
- package/termux_tts/engine_multilingual.py +374 -0
- package/termux_tts/engine_sherpa.py +2 -1
- package/termux_tts/engine_sherpa_capi.py +491 -0
- package/termux_tts/hardware.py +34 -1
- package/termux_tts/installer.py +213 -11
- package/termux_tts/script_classifier.py +253 -0
- package/termux_tts/tokenizer.py +3 -1
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to this project will be documented in this file.
|
|
4
4
|
|
|
5
|
+
## [1.4.4] - 2026-09-14
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
- **Multilingual Neural Orchestrator**: Integrated `MultilingualNeuralEngine` capable of dynamic cross-language code-switching across 9 official languages (ko, en, ja, zh, hi, ru, es, fr, de).
|
|
9
|
+
- **Resident C-API Acceleration**: Implemented `SherpaResidentManager` and `SherpaCapiSession` for in-memory model caching via direct C-API bindings, achieving sub-0.18x RTF and eliminating subprocess invocation overhead.
|
|
10
|
+
- **Zero-Config CLI Ergonomics**: Promoted speech synthesis (`synth`) as the default top-level subcommand; allows direct positional text (`termux-tts "text" --play`); eliminated necessity of `-e multilingual` flag.
|
|
11
|
+
- **Universal Unicode Script Classifier**: Built `MultilingualTokenizer` covering Hangul, Latin, Devanagari (Hindi), Cyrillic (Russian), CJK, and Arabic scripts with 50ms smooth pause padding.
|
|
12
|
+
- **On-Demand Model Provisioning & Self-Healing**: Enhanced `termux-tts install` to provision English/Korean by default, with on-demand flags (`--models hi/ja/ru/zh/all`) and actionable English guidance with absolute paths for uninstalled models.
|
|
13
|
+
- **Defensive Path Resolution**: Hardened path parsing across `AudioBuffer.save()` and CLI entrypoints with automatic directory creation and `Path.expanduser()` resolution.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## [1.4.3] - 2026-09-07
|
|
18
|
+
|
|
19
|
+
### Added
|
|
20
|
+
- **Multi-Tier Dynamic Candidate Resolution**: Replaced single hardcoded `v1.0.0-vulkan` URL in `installer.py` with 4-tier candidate resolution (`TERMUX_TTS_RELEASE_TAG`, `v{__version__}`, `releases/latest/download`, `uno-km/ameva-runtime` SSOT fallback, and companion fallback).
|
|
21
|
+
- **Dual-Path Installation**: Installs and links precompiled Vulkan binary `sherpa-ncnn-offline-tts` to both `~/.local/bin` and `$PREFIX/bin` for instant global PATH resolution.
|
|
22
|
+
- **Dynamic SSOT User-Agent**: Injects dynamic package version into download requests.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
5
26
|
## [1.4.2] - 2026-09-07
|
|
6
27
|
|
|
7
28
|
### Added
|
package/README.md
CHANGED
|
@@ -1,438 +1,138 @@
|
|
|
1
|
-
# Termux-TTS
|
|
1
|
+
# Termux-TTS
|
|
2
2
|
|
|
3
3
|
[](https://pypi.org/project/termux-tts/)
|
|
4
4
|
[](https://pypi.org/project/termux-tts/)
|
|
5
5
|
[](https://www.npmjs.com/package/termux-tts)
|
|
6
6
|
[](https://github.com/uno-km/termux-tts)
|
|
7
|
-
[](https://www.vulkan.org/)
|
|
8
7
|
|
|
9
|
-
>
|
|
8
|
+
> Production-Grade 4-Tier On-Device Speech Synthesis Framework (Multilingual Neural Orchestrator, Zero-Dependency DSP Formant & C++ Acceleration)
|
|
10
9
|
|
|
11
10
|
---
|
|
12
11
|
|
|
13
|
-
##
|
|
12
|
+
## Architecture & Overview
|
|
14
13
|
|
|
15
|
-
Termux-TTS is
|
|
14
|
+
Termux-TTS is an enterprise-grade, on-device text-to-speech framework optimized for mobile edge hardware and Android Termux environments. Built to eliminate heavy dependency stacks and fragile driver behaviors, it features a resilient 4-Tier architecture:
|
|
16
15
|
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
16
|
+
- **Multilingual Neural Orchestrator**: Dynamic cross-language code-switching mesh covering 9 official languages (Korean, English, Japanese, Chinese, Hindi, Russian, Spanish, French, German) with Unicode script tokenization and 50ms context-aware silence padding.
|
|
17
|
+
- **Resident C-API In-Memory Engine**: Direct C-API memory residency with `SherpaResidentManager` achieving sub-0.18x real-time factor with ARM NEON SIMD acceleration and zero subprocess lag.
|
|
18
|
+
- **Tier 1: Zero-Dependency Parametric DSP Formant Vocoder**: 0MB disk footprint, Rosenberg glottal pulse formulation, and 5-band biquad formant filters providing deterministic speech synthesis in under 50 milliseconds (RTF 0.013x).
|
|
19
|
+
- **Tier 2: Android System Native Voice Bridge**: Direct IPC integration to physical Samsung and Google speech engines via the Termux-API service layer.
|
|
20
|
+
- **Tier 3: Subprocess-Isolated Sherpa C++ CPU Engine**: Subprocess-isolated VITS acoustic modeling on ARM64 NEON with memory leak protection.
|
|
21
|
+
- **Tier 4: Pure Vulkan GPU Hardware Neural Engine**: High-performance GPU tensor synthesis via precompiled native C++ binaries (`sherpa-ncnn-offline-tts-vulkan`) running high-resolution studio models (`vits-piper-en_US-lessac-high-fp16`) with zero silent CPU fallback.
|
|
23
22
|
|
|
24
|
-
|
|
25
|
-
Install the core package from PyPI via `pip`:
|
|
26
|
-
```bash
|
|
27
|
-
pip install --upgrade pip
|
|
28
|
-
pip install termux-tts
|
|
29
|
-
```
|
|
23
|
+
---
|
|
30
24
|
|
|
31
|
-
|
|
32
|
-
```bash
|
|
33
|
-
pip install "termux-tts[neural,dev]"
|
|
34
|
-
```
|
|
25
|
+
## Empirical Hardware Benchmarks (Physical Devices)
|
|
35
26
|
|
|
36
|
-
|
|
37
|
-
Install globally or locally inside your Node.js application via `npm`:
|
|
38
|
-
```bash
|
|
39
|
-
# Global CLI installation
|
|
40
|
-
npm install -g termux-tts
|
|
27
|
+
Measurements gathered on physical Android 16 hardware running Termux ARM64:
|
|
41
28
|
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
29
|
+
| Target Device | Hardware Architecture | Synthesis Engine | Model Profile | Audio Length | Synthesis Time | Real-Time Factor (RTF) | Status |
|
|
30
|
+
| :--- | :--- | :--- | :--- | :--- | :---: | :---: | :---: |
|
|
31
|
+
| **Galaxy A53** | Exynos 1280 / ARM64 NEON | Multilingual Neural C-API | `KSS + Lessac` (Bilingual) | 10.50 s | **1.82 s** | **0.173x** | Validated (5.78x RT) |
|
|
32
|
+
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `lessac-high-fp16` | 6.70 s | **6.65 s** | **0.993x** | Validated |
|
|
33
|
+
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `lessac-medium` | 4.59 s | **1.21 s** | **0.264x** | Validated (3.79x RT) |
|
|
34
|
+
| **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU Neural | `lessac-medium` | 4.52 s | **5.18 s** | **1.146x** | Validated |
|
|
35
|
+
| **ARM64 CPU** | Cortex-A78 / A55 | Parametric DSP | 5-Band Biquad | 4.15 s | **0.054 s** | **0.0130x** | Validated (76x RT) |
|
|
45
36
|
|
|
46
37
|
---
|
|
47
38
|
|
|
48
|
-
##
|
|
49
|
-
|
|
50
|
-
To unlock raw mobile GPU compute via Vulkan SPIR-V compute pipelines on Qualcomm Adreno or ARM Mali silicon, pair `termux-tts` with the unified `@ameva/runtime` hardware acceleration layer and run the automated 1-click provisioning tool.
|
|
51
|
-
|
|
52
|
-
### 2.1 Unified Installation Command
|
|
53
|
-
Install both the speech engine and the hardware acceleration runtime simultaneously:
|
|
39
|
+
## Installation & Automated Provisioning
|
|
54
40
|
|
|
41
|
+
### 1. Package Installation
|
|
55
42
|
```bash
|
|
56
|
-
# Python
|
|
57
|
-
pip install termux-tts
|
|
43
|
+
# Python SDK & CLI
|
|
44
|
+
pip install termux-tts
|
|
58
45
|
|
|
59
|
-
# Node.js /
|
|
60
|
-
npm install
|
|
46
|
+
# Node.js / TypeScript
|
|
47
|
+
npm install termux-tts
|
|
61
48
|
```
|
|
62
49
|
|
|
63
|
-
### 2.
|
|
64
|
-
|
|
65
|
-
|
|
50
|
+
### 2. Automated Model Provisioning
|
|
51
|
+
Automate downloading and linking precompiled models to the official immutable path (`/data/data/com.termux/files/home/models/tts/`):
|
|
66
52
|
```bash
|
|
67
|
-
#
|
|
68
|
-
termux-tts install
|
|
53
|
+
# Default lightweight installation (Korean KSS + English Lessac)
|
|
54
|
+
termux-tts install
|
|
69
55
|
|
|
70
|
-
#
|
|
71
|
-
termux-tts install --
|
|
56
|
+
# On-demand provision specific language models:
|
|
57
|
+
termux-tts install --models hi # Hindi (Piper Swara)
|
|
58
|
+
termux-tts install --models ja # Japanese (Piper Hina)
|
|
59
|
+
termux-tts install --models ru # Russian (Piper Dmitri)
|
|
60
|
+
termux-tts install --models zh # Chinese (AISHELL3)
|
|
61
|
+
termux-tts install --models all # All 9 official languages
|
|
72
62
|
|
|
73
|
-
#
|
|
74
|
-
termux-tts install --tier high
|
|
75
|
-
```
|
|
76
|
-
|
|
77
|
-
Once installed, verify full Vulkan GPU compute availability:
|
|
78
|
-
```bash
|
|
79
|
-
termux-tts doctor
|
|
63
|
+
# Provision Studio Vulkan GPU engine:
|
|
64
|
+
termux-tts install --tier high
|
|
80
65
|
```
|
|
81
66
|
|
|
82
67
|
---
|
|
83
68
|
|
|
84
|
-
##
|
|
85
|
-
|
|
86
|
-
Termux-TTS provides intuitive interfaces across CLI, Python, and Node.js.
|
|
69
|
+
## Quickstart
|
|
87
70
|
|
|
88
|
-
###
|
|
71
|
+
### Global Command-Line Interface (Zero-Config)
|
|
89
72
|
```bash
|
|
90
|
-
# 1.
|
|
91
|
-
termux-tts
|
|
73
|
+
# 1. Zero-config synthesis with physical speaker playback
|
|
74
|
+
termux-tts "Hello 방가방가 나는 parrot 이라고 해." -o ~/out.wav --play
|
|
75
|
+
|
|
76
|
+
# 2. Flag-based syntax
|
|
77
|
+
termux-tts -t "Multilingual neural speech synthesis on device." --play
|
|
92
78
|
|
|
93
|
-
#
|
|
94
|
-
termux-tts
|
|
79
|
+
# 3. Force language pinning
|
|
80
|
+
termux-tts -l en -t "Pure English text output." --play
|
|
81
|
+
termux-tts -l ko -t "한국어 단독 신경망 음성 합성." --play
|
|
95
82
|
|
|
96
|
-
#
|
|
97
|
-
termux-tts
|
|
83
|
+
# 4. Instant DSP Formant synthesis (0MB footprint)
|
|
84
|
+
termux-tts synth -e dsp -t "Zero dependency DSP synthesis." -o dsp.wav
|
|
98
85
|
|
|
99
|
-
#
|
|
100
|
-
termux-tts
|
|
86
|
+
# 5. Direct Android system native speaker broadcast
|
|
87
|
+
termux-tts speak -t "Hardware speaker broadcast via Android service."
|
|
88
|
+
|
|
89
|
+
# 6. Hardware diagnostics
|
|
90
|
+
termux-tts doctor
|
|
101
91
|
```
|
|
102
92
|
|
|
103
|
-
###
|
|
93
|
+
### Python SDK
|
|
104
94
|
```python
|
|
105
95
|
import termux_tts as tts
|
|
106
96
|
|
|
107
|
-
#
|
|
108
|
-
with tts.load(
|
|
109
|
-
|
|
110
|
-
"
|
|
111
|
-
output="
|
|
112
|
-
speed=1.0
|
|
97
|
+
# 1. Zero-Config Multilingual Neural Synthesis (Korean + English Code-Switching)
|
|
98
|
+
with tts.load() as engine:
|
|
99
|
+
result = engine.synthesize(
|
|
100
|
+
"Hello 방가방가 키키키키 나는 parrot 이라고 해.",
|
|
101
|
+
output="multilingual.wav"
|
|
113
102
|
)
|
|
114
|
-
print(f"
|
|
103
|
+
print(f"Generated {result.duration_sec:.2f}s in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
|
|
115
104
|
|
|
116
|
-
#
|
|
117
|
-
with tts.load(engine="
|
|
118
|
-
|
|
119
|
-
print(f"
|
|
105
|
+
# 2. Pure Vulkan GPU Neural Synthesis (Studio Tier)
|
|
106
|
+
with tts.load(engine="vulkan", model_tier="high") as engine:
|
|
107
|
+
result = engine.synthesize("Pure Vulkan neural execution on mobile.", output="speech.wav")
|
|
108
|
+
print(f"Synthesized in {result.elapsed_ms:.1f}ms (RTF: {result.rtf:.4f}x)")
|
|
120
109
|
|
|
121
|
-
#
|
|
122
|
-
with tts.load(engine="
|
|
123
|
-
engine.
|
|
110
|
+
# 3. Zero-Dependency DSP Formant Synthesis
|
|
111
|
+
with tts.load(engine="dsp", preset="balanced") as engine:
|
|
112
|
+
result = engine.synthesize("Instant speech without model downloads.", output="dsp.wav")
|
|
113
|
+
print(f"DSP Latency: {result.elapsed_ms:.1f}ms")
|
|
124
114
|
```
|
|
125
115
|
|
|
126
|
-
###
|
|
116
|
+
### Node.js / TypeScript
|
|
127
117
|
```typescript
|
|
128
118
|
import * as tts from 'termux-tts';
|
|
129
119
|
|
|
130
120
|
async function main() {
|
|
131
|
-
|
|
132
|
-
const
|
|
133
|
-
|
|
134
|
-
tier: 'high',
|
|
135
|
-
language: 'en'
|
|
136
|
-
});
|
|
137
|
-
|
|
138
|
-
const result = await engine.synthesize(
|
|
139
|
-
"Synthesizing high-resolution neural speech in Node.js on Termux.",
|
|
140
|
-
{ output: 'node_output.wav', speed: 1.0 }
|
|
141
|
-
);
|
|
142
|
-
|
|
143
|
-
console.log(`Synthesized in ${result.elapsedMs}ms | RTF: ${result.rtf}x`);
|
|
121
|
+
const engine = tts.load({ engine: 'vulkan', tier: 'high' });
|
|
122
|
+
const res = await engine.synthesize("High performance speech synthesis.", { output: "speech.wav" });
|
|
123
|
+
console.log(`Synthesized in ${res.elapsedMs}ms`);
|
|
144
124
|
}
|
|
145
|
-
|
|
146
|
-
main().catch(console.error);
|
|
147
|
-
```
|
|
148
|
-
|
|
149
|
-
---
|
|
150
|
-
|
|
151
|
-
## 4. Advanced Usage & 4-Tier Speech Architecture
|
|
152
|
-
|
|
153
|
-
Termux-TTS features a robust 4-tier synthesis architecture engineered to provide the optimal balance between acoustic fidelity, memory consumption, and compute latency.
|
|
154
|
-
|
|
155
|
-
### 4.1 Architecture Tier Breakdown
|
|
156
|
-
1. **Tier 1: Parametric DSP Formant Vocoder (`engine="dsp"`)**:
|
|
157
|
-
- Zero-dependency Rosenberg glottal source pulse generator paired with a 5-band second-order biquad formant resonator filter bank.
|
|
158
|
-
- 0MB disk footprint, deterministic latency under 50ms (RTF ~0.013x on ARM Cortex-A78).
|
|
159
|
-
- Ideal for embedded alerts, battery-saving modes, and fail-safe recovery.
|
|
160
|
-
2. **Tier 2: Android System Native Voice Bridge (`engine="native"`)**:
|
|
161
|
-
- Direct IPC connection to Android `TextToSpeech` service via `termux-api` and Unix domain sockets.
|
|
162
|
-
- Zero CPU inference overhead; delegates synthesis and playback to Samsung Voice Engine or Google Speech Services.
|
|
163
|
-
3. **Tier 3: Subprocess-Isolated Sherpa C++ Engine (`engine="neural"`)**:
|
|
164
|
-
- CPU-based VITS neural acoustic synthesis with multi-threaded ARM NEON SIMD vectorization.
|
|
165
|
-
- Subprocess process-isolation prevents memory fragmentation and native memory leaks during multi-hour continuous execution.
|
|
166
|
-
4. **Tier 4: Pure Vulkan GPU Hardware Neural Engine (`engine="vulkan"`)**:
|
|
167
|
-
- End-to-end GPU compute shader pipeline via `sherpa-ncnn` with zero silent CPU fallback.
|
|
168
|
-
- Evaluates high-resolution FP16 neural models (`vits-piper-en_US-lessac-high-fp16`) natively on mobile GPU silicon.
|
|
169
|
-
|
|
170
|
-
### 4.2 Emotional & Expressive Conversational Modulation
|
|
171
|
-
The Expressive Engine (`engine="expressive"`) injects natural acoustic non-verbal vocalizations directly into neural synthesis:
|
|
172
|
-
- `[sigh]` / `[한숨]`: Organic aspiration decay and exhalation acoustic wave.
|
|
173
|
-
- `[laugh]` / `[웃음]`: Rhythmic 6.5 Hz glottal laughter bursts with vocal fold resonance.
|
|
174
|
-
- `[breath]` / `[호흡]`: Soft physiological inhalation pause.
|
|
175
|
-
- `[pause]` / `[쉼]`: Contextual silence spacing.
|
|
176
|
-
|
|
177
|
-
```python
|
|
178
|
-
import termux_tts as tts
|
|
179
|
-
|
|
180
|
-
with tts.load(engine="expressive", language="en") as engine:
|
|
181
|
-
script = (
|
|
182
|
-
"Good morning! [breath] We have successfully deployed the system. "
|
|
183
|
-
"[laugh] It took all night, [sigh] but everything is operating smoothly now."
|
|
184
|
-
)
|
|
185
|
-
res = engine.synthesize(script, output="conversational.wav")
|
|
186
|
-
print(f"Expressive tags processed: {res.expressive_tags_detected}")
|
|
187
|
-
```
|
|
188
|
-
|
|
189
|
-
### 4.3 Streaming & Real-Time Audio Buffer Manipulation
|
|
190
|
-
Inspect and manipulate raw floating-point and 16-bit PCM audio buffers directly in memory before writing to disk:
|
|
191
|
-
```python
|
|
192
|
-
import termux_tts as tts
|
|
193
|
-
|
|
194
|
-
with tts.load(engine="dsp") as engine:
|
|
195
|
-
res = engine.synthesize("Buffer streaming test.")
|
|
196
|
-
audio_buf = res.audio_buffer
|
|
197
|
-
|
|
198
|
-
# Access raw PCM sample array
|
|
199
|
-
samples = audio_buf.samples # np.ndarray (float32, normalized [-1.0, 1.0])
|
|
200
|
-
raw_bytes = audio_buf.to_wav_bytes() # RIFF WAV binary stream
|
|
201
|
-
print(f"Sample count: {len(samples)}, Duration: {audio_buf.duration_seconds:.3f}s")
|
|
125
|
+
main();
|
|
202
126
|
```
|
|
203
127
|
|
|
204
128
|
---
|
|
205
129
|
|
|
206
|
-
##
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
| Option Flag | Argument Type | Default | Description |
|
|
211
|
-
| :--- | :--- | :--- | :--- |
|
|
212
|
-
| `-t`, `--text` | `string` | *(Required)* | Input text or SSML-tagged phrase to synthesize. |
|
|
213
|
-
| `-o`, `--output` | `path` | `output.wav` | Destination filesystem path for generated RIFF WAV file. |
|
|
214
|
-
| `-l`, `--lang` | `string` | `ko` | Target language code (`en`, `ko`). |
|
|
215
|
-
| `-e`, `--engine` | `enum` | `auto` | Engine backend tier: `auto`, `vulkan`, `dsp`, `native`, `neural`, `expressive`. |
|
|
216
|
-
| `--tier` | `enum` | `high` | Model resolution profile: `high` (Studio FP16, 57MB), `medium` (Balanced, 25MB). |
|
|
217
|
-
| `-p`, `--preset` | `enum` | `balanced` | DSP vocoder profile: `fast`, `balanced`, `expressive`, `ultra`. |
|
|
218
|
-
| `-d`, `--device` | `enum` | `auto` | Compute execution device: `auto`, `gpu`, `vulkan`, `cpu`. |
|
|
219
|
-
| `-s`, `--speed` | `float` | `1.0` | Speech cadence multiplier (`0.5` to `2.0`). |
|
|
220
|
-
| `--threads` | `int` | `4` | Worker threads for CPU ARM NEON SIMD compute (1 for GPU). |
|
|
221
|
-
| `--volume` | `int` | `None` | Android media volume level (`1` to `15`). |
|
|
222
|
-
| `--play` | `flag` | `False` | Automatically dispatch audio to physical speaker upon synthesis completion. |
|
|
223
|
-
|
|
224
|
-
### 5.2 Python SDK `load()` Parameters
|
|
225
|
-
|
|
226
|
-
| Parameter | Type | Default | Description |
|
|
227
|
-
| :--- | :--- | :--- | :--- |
|
|
228
|
-
| `engine` | `str` | `"auto"` | Selects target engine: `"vulkan"`, `"dsp"`, `"native"`, `"neural"`, `"expressive"`. |
|
|
229
|
-
| `tier` | `str` | `"high"` | Specifies neural weight tier (`"high"`, `"medium"`). |
|
|
230
|
-
| `language` | `str` | `"ko"` | Phonemizer and lexicon locale code. |
|
|
231
|
-
| `preset` | `str` | `"balanced"` | Parametric DSP quality level (`"fast"`, `"balanced"`, `"expressive"`, `"ultra"`). |
|
|
232
|
-
| `device` | `str` | `"auto"` | Compute target device (`"gpu"`, `"vulkan"`, `"cpu"`, `"auto"`). |
|
|
233
|
-
| `threads` | `int` | `4` | CPU concurrency worker count. |
|
|
234
|
-
| `model` | `str` | `None` | Optional explicit path to custom VITS or NCNN model directory. |
|
|
235
|
-
| `sample_rate`| `int` | `22050` | Audio sampling frequency (Hz). |
|
|
236
|
-
|
|
237
|
-
---
|
|
238
|
-
|
|
239
|
-
## 6. Production Code Examples & Diagnostics
|
|
240
|
-
|
|
241
|
-
### 6.1 Enterprise Batch Speech Pipeline with Fallback Assurance
|
|
242
|
-
```python
|
|
243
|
-
import os
|
|
244
|
-
import termux_tts as tts
|
|
245
|
-
from termux_tts.exceptions import VulkanInitializationError, TTSModelLoadError
|
|
246
|
-
|
|
247
|
-
scripts = [
|
|
248
|
-
"Alert: Thermal gradient within nominal thresholds.",
|
|
249
|
-
"System diagnostics passed all 12 validation gates.",
|
|
250
|
-
"Unattended background daemon active on port 8080."
|
|
251
|
-
]
|
|
252
|
-
|
|
253
|
-
def synthesize_batch(items, out_dir="dist_audio"):
|
|
254
|
-
os.makedirs(out_dir, exist_ok=True)
|
|
255
|
-
|
|
256
|
-
# Attempt primary Tier 4 Vulkan GPU engine with Fail-Safe fallback to Tier 1 DSP
|
|
257
|
-
try:
|
|
258
|
-
engine = tts.load(engine="vulkan", tier="high")
|
|
259
|
-
print("[INFO] Initialized Tier 4 Vulkan GPU Neural Engine.")
|
|
260
|
-
except (VulkanInitializationError, TTSModelLoadError) as exc:
|
|
261
|
-
print(f"[WARN] Hardware acceleration unavailable ({exc}). Falling back to Tier 1 DSP.")
|
|
262
|
-
engine = tts.load(engine="dsp", preset="balanced")
|
|
263
|
-
|
|
264
|
-
with engine:
|
|
265
|
-
for idx, text in enumerate(items):
|
|
266
|
-
out_file = os.path.join(out_dir, f"notice_{idx:02d}.wav")
|
|
267
|
-
res = engine.synthesize(text, output=out_file)
|
|
268
|
-
print(f"[{idx+1}/{len(items)}] Generated '{out_file}' | Backend: {res.backend} | RTF: {res.rtf:.4f}x")
|
|
269
|
-
|
|
270
|
-
if __name__ == "__main__":
|
|
271
|
-
synthesize_batch(scripts)
|
|
272
|
-
```
|
|
273
|
-
|
|
274
|
-
### 6.2 12-Stage Hardware Diagnostic Validation
|
|
275
|
-
Run programmatic health audits to verify Vulkan compute queues, shared memory bindings, and driver integrity:
|
|
276
|
-
```python
|
|
277
|
-
from termux_tts.engine import doctor
|
|
278
|
-
|
|
279
|
-
report = doctor()
|
|
280
|
-
print("Hardware Diagnostic Status:", report["status"])
|
|
281
|
-
print("Passed Verification Stages:", report["passed_stages"])
|
|
282
|
-
print("Recommended Compute Backend:", report["recommended_backend"])
|
|
283
|
-
```
|
|
284
|
-
|
|
285
|
-
---
|
|
286
|
-
|
|
287
|
-
## 7. Real-World Outputs & Empirical Hardware Benchmarks
|
|
288
|
-
|
|
289
|
-
### 7.1 Empirical Physical Device Benchmarks
|
|
290
|
-
All metrics were gathered directly on physical mobile hardware running Android 16 / Termux ARM64:
|
|
291
|
-
|
|
292
|
-
| Device Model | Processor Architecture | Synthesis Engine | Model Profile | Audio Length | Inference Time | Real-Time Factor (RTF) | Memory Footprint |
|
|
293
|
-
| :--- | :--- | :--- | :--- | :---: | :---: | :---: | :---: |
|
|
294
|
-
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `lessac-high-fp16` | 6.70 s | **6.65 s** | **0.993x** | 68 MB |
|
|
295
|
-
| **Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | Vulkan GPU Neural | `amy-medium` | 4.59 s | **1.21 s** | **0.264x** | 38 MB |
|
|
296
|
-
| **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU Neural | `amy-medium` | 4.52 s | **5.18 s** | **1.146x** | 42 MB |
|
|
297
|
-
| **Galaxy A35** | Exynos 1380 / Mali-G68 MP5 | Vulkan GPU Neural | `lessac-high-fp16` | 6.73 s | **34.33 s** | **5.098x** | 72 MB |
|
|
298
|
-
| **ARM64 CPU** | Cortex-A78 / A55 (Quad-Core) | Parametric DSP Vocoder | 5-Band Biquad | 4.15 s | **0.054 s** | **0.0130x** | **0 MB** |
|
|
299
|
-
|
|
300
|
-
> **Real-Time Factor (RTF) Definition**: $\text{RTF} = \frac{\text{Synthesis Latency (Seconds)}}{\text{Generated Audio Duration (Seconds)}}$.
|
|
301
|
-
> An RTF under `1.0x` indicates faster-than-realtime synthesis suitable for live interactive voice applications.
|
|
302
|
-
|
|
303
|
-
### 7.2 Verified Audio Samples
|
|
304
|
-
Reference audio samples generated directly on-device are included in the repository:
|
|
305
|
-
- **Expressive Emotional Output**: [`docs/assets/samples/expressive_demo.wav`](https://github.com/uno-km/termux-tts/blob/main/docs/assets/samples/expressive_demo.wav) — Demonstrates natural aspiration sighs and laughter tags.
|
|
306
|
-
- **Parametric DSP Output**: [`docs/assets/samples/dsp_test.wav`](https://github.com/uno-km/termux-tts/blob/main/docs/assets/samples/dsp_test.wav) — Demonstrates 0MB instant formant synthesis.
|
|
307
|
-
|
|
308
|
-
---
|
|
309
|
-
|
|
310
|
-
## 8. GPU Interconnect Architecture & Compatibility
|
|
311
|
-
|
|
312
|
-
```mermaid
|
|
313
|
-
flowchart LR
|
|
314
|
-
A["Termux-TTS Application Layer"] --> B["AMEVA Hardware Gateway"]
|
|
315
|
-
B --> C["/system/lib64/libvulkan.so"]
|
|
316
|
-
C --> D{"SoC GPU Silicon"}
|
|
317
|
-
D -->|"Adreno 7xx / 8xx (Full SPIR-V FP16)"| E["Qualcomm Snapdragon"]
|
|
318
|
-
D -->|"Mali Bifrost / Valhall (Driver Pipelined)"| F["ARM Mali / Exynos"]
|
|
319
|
-
E --> G["High-Throughput Shader Core (RTF < 0.3x)"]
|
|
320
|
-
F --> H["Balanced Execution (Medium Tier Recommended)"]
|
|
321
|
-
```
|
|
322
|
-
|
|
323
|
-
### 8.1 Vulkan Compute Shader Pipeline
|
|
324
|
-
Termux-TTS interfaces directly with `/system/lib64/libvulkan.so` via SPIR-V compute shaders compiled in `sherpa-ncnn`. Tensor matrix multiplications for VITS encoder, duration predictor, and inverse coupling flows execute directly on GPU compute units.
|
|
325
|
-
|
|
326
|
-
### 8.2 Silicon Compatibility Matrix
|
|
327
|
-
- **Qualcomm Snapdragon (Adreno 6xx, 7xx, 8xx)**:
|
|
328
|
-
- **Status: Tier-1 Full Support**. Hardware FP16 arithmetic instructions, high sub-group sizes, and low dispatch latency deliver real-time factor performance as low as `0.264x`.
|
|
329
|
-
- **Samsung Exynos / MediaTek Dimensity (ARM Mali-Gxx / Immortalis)**:
|
|
330
|
-
- **Status: Supported (Medium Tier Recommended)**. Works out of the box. Due to Mali driver SPIR-V shader compilation overhead, `--tier medium` (`amy-medium`) is recommended for real-time responsiveness.
|
|
331
|
-
- **Strict Zero-Silent-Fallback**:
|
|
332
|
-
- If `--engine vulkan` or `--gpu` is specified and no Vulkan driver or compatible hardware is available, Termux-TTS raises `VulkanInitializationError` immediately rather than secretly degrading to CPU execution.
|
|
333
|
-
|
|
334
|
-
---
|
|
335
|
-
|
|
336
|
-
## 8-1. CPU vs. GPU Performance & Thermal Trade-offs
|
|
337
|
-
|
|
338
|
-
| Evaluation Metric | CPU Synthesis (ARM Cortex-A78) | Vulkan GPU Neural (Adreno 830) | Parametric DSP (0MB) |
|
|
339
|
-
| :--- | :--- | :--- | :--- |
|
|
340
|
-
| **Real-Time Factor (Medium)** | ~0.85x – 1.10x | **0.264x** (3.5x Faster) | **0.013x** (70x Faster) |
|
|
341
|
-
| **Real-Time Factor (Studio High)** | ~3.80x – 5.20x | **0.993x** (Sub-realtime) | N/A (Formant Only) |
|
|
342
|
-
| **First-Token Latency (TTFA)** | ~450 ms | ~180 ms | **< 15 ms** |
|
|
343
|
-
| **CPU Big-Core Utilization** | 100% across 4 cores | < 15% (Driver Dispatch) | Single Core ~8% |
|
|
344
|
-
| **Thermal Dissipation** | High (Thermal Throttling at ~3min) | Low to Moderate | Negligible |
|
|
345
|
-
| **Memory Allocation** | ~85 MB Heap | ~38 MB (GPU VRAM Mapped) | **0 MB Disk / < 2MB RAM** |
|
|
346
|
-
|
|
347
|
-
Offloading neural acoustic calculations to the Vulkan GPU protects CPU big cores from thermal throttling during prolonged text reading, maintaining consistent interactive responsiveness across Android background services.
|
|
348
|
-
|
|
349
|
-
---
|
|
350
|
-
|
|
351
|
-
## 9. Hardware Requirements & Operational Limits
|
|
352
|
-
|
|
353
|
-
### 9.1 Hardware Specifications
|
|
354
|
-
|
|
355
|
-
| Specification Metric | Minimum Requirements | Recommended Production Spec |
|
|
356
|
-
| :--- | :--- | :--- |
|
|
357
|
-
| **Operating System** | Android 9.0+ (API level 28+) / Linux 5.4+ | Android 12.0+ (API level 31+) |
|
|
358
|
-
| **Architecture** | ARM64 (aarch64) or x86_64 | ARM64-v8a / v9a |
|
|
359
|
-
| **System RAM** | 2 GB Total (DSP Tier: 512 MB) | 4 GB+ Unified RAM |
|
|
360
|
-
| **Storage Footprint** | 10 MB (DSP Only) / 80 MB (Neural) | 250 MB Free Flash Storage |
|
|
361
|
-
| **GPU Subsystem** | Vulkan 1.1 Conforming Mobile Driver | Qualcomm Adreno 660 / 730 / 830 or Mali-G78+ |
|
|
362
|
-
|
|
363
|
-
### 9.2 Known Operational Limits
|
|
364
|
-
- **32-Bit ARM (armeabi-v7a)**: Not supported for Vulkan GPU neural compute. Use Tier 1 DSP vocoder for legacy 32-bit hardware.
|
|
365
|
-
- **Headless SSH Environments**: Audio playback (`--play`) requires Termux-API or PulseAudio daemon running. To output directly without audio hardware, synthesize to `.wav` file.
|
|
366
|
-
|
|
367
|
-
---
|
|
368
|
-
|
|
369
|
-
## 10. 24/7 Unattended Background Execution Guide
|
|
370
|
-
|
|
371
|
-
Android aggressively kills background user-space processes running inside Termux unless battery and process monitor policies are explicitly configured. Follow these three stages to ensure uninterrupted 24/7 autonomous speech services:
|
|
372
|
-
|
|
373
|
-
### 10.1 Stage 1: Termux Wake-Lock
|
|
374
|
-
Prevent the Android kernel from entering deep CPU sleep states:
|
|
375
|
-
```bash
|
|
376
|
-
# Acquire persistent CPU wake-lock
|
|
377
|
-
termux-wake-lock
|
|
378
|
-
```
|
|
379
|
-
|
|
380
|
-
### 10.2 Stage 2: Android OS GUI Settings
|
|
381
|
-
1. Navigate to **Android Settings > Apps > Termux > Battery**.
|
|
382
|
-
2. Select **Unrestricted** (Disable battery optimization).
|
|
383
|
-
3. Under **Permissions**, grant **Notifications** and **Display over other apps** (if applicable).
|
|
384
|
-
|
|
385
|
-
### 10.3 Stage 3: ADB Phantom Process Killer Mitigation (Android 12+)
|
|
386
|
-
Android 12 introduced the Phantom Process Killer, which terminates child processes exceeding 32 instances or high CPU thresholds. Execute the following commands via PC ADB or wireless debugging:
|
|
387
|
-
|
|
388
|
-
```bash
|
|
389
|
-
# Disable Android Phantom Process Killer
|
|
390
|
-
adb shell device_config put activity_manager max_phantom_processes 2147483647
|
|
391
|
-
adb shell settings put global settings_enable_monitor_phantom_procs false
|
|
392
|
-
|
|
393
|
-
# Verify configuration
|
|
394
|
-
adb shell settings get global settings_enable_monitor_phantom_procs
|
|
395
|
-
# Expected output: false
|
|
396
|
-
```
|
|
397
|
-
|
|
398
|
-
---
|
|
399
|
-
|
|
400
|
-
## 11. Open Source License
|
|
401
|
-
|
|
402
|
-
Termux-TTS is open-sourced under the **Apache License, Version 2.0**.
|
|
403
|
-
|
|
404
|
-
```text
|
|
405
|
-
Copyright 2026 Eunho Kim (@uno-km) & AMEVA Open-Source Foundation.
|
|
406
|
-
|
|
407
|
-
Licensed under the Apache License, Version 2.0 (the "License");
|
|
408
|
-
you may not use this file except in compliance with the License.
|
|
409
|
-
You may obtain a copy of the License at
|
|
410
|
-
|
|
411
|
-
http://www.apache.org/licenses/LICENSE-2.0
|
|
412
|
-
|
|
413
|
-
Unless required by applicable law or agreed to in writing, software
|
|
414
|
-
distributed under the License is distributed on an "AS IS" BASIS,
|
|
415
|
-
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
416
|
-
See the License for the specific language governing permissions and
|
|
417
|
-
limitations under the License.
|
|
418
|
-
```
|
|
419
|
-
|
|
420
|
-
### Key Licensing Permissions & Terms:
|
|
421
|
-
- **Commercial Use**: Permitted without royalty or proprietary source disclosure.
|
|
422
|
-
- **Modification & Distribution**: Permitted provided that modified files carry prominent notices.
|
|
423
|
-
- **Patent Grant**: Express grant of patent rights from contributors.
|
|
424
|
-
- **Trademark**: Does not grant permission to use project trademark names without prior written consent.
|
|
425
|
-
- **No Warranty & Limitation of Liability**: Software is provided strictly on an "AS IS" basis.
|
|
426
|
-
|
|
427
|
-
---
|
|
428
|
-
|
|
429
|
-
## 12. SEO Technical Keywords & Metadata
|
|
430
|
-
|
|
431
|
-
`tts`, `text-to-speech`, `vulkan`, `vulkan-compute`, `vits`, `piper-tts`, `sherpa-onnx`, `sherpa-ncnn`, `termux`, `android`, `on-device-ai`, `speech-synthesis`, `edge-ai`, `formant-synthesis`, `vocoder`, `mobile-ai`, `adreno`, `mali-gpu`, `dsp`, `rosenberg-glottal`, `biquad-filter`, `expressive-speech`, `voice-cloning`, `audio-generation`, `ncnn`, `arm64`, `snapdragon`, `exynos`, `real-time-factor`, `low-latency`, `zero-dependency`, `voice-assistant`, `headless-audio`, `embedded-systems`
|
|
130
|
+
## Official Documentation & Benchmarks
|
|
131
|
+
- [Official Architecture & API Reference](https://uno-km.vercel.app/lib/tts/)
|
|
132
|
+
- [Ecosystem Metrics & Registry Stats](https://uno-km.vercel.app/foundation/metrics)
|
|
133
|
+
- [AMEVA Open-Source Foundation Portal](https://uno-km.vercel.app/foundation/index.html)
|
|
432
134
|
|
|
433
135
|
---
|
|
434
136
|
|
|
435
|
-
##
|
|
436
|
-
-
|
|
437
|
-
- **GitHub Repository**: [https://github.com/uno-km/termux-tts](https://github.com/uno-km/termux-tts)
|
|
438
|
-
- **AMEVA Foundation Portal**: [https://uno-km.vercel.app/foundation/index.html](https://uno-km.vercel.app/foundation/index.html)
|
|
137
|
+
## License
|
|
138
|
+
Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)).
|