termux-stt 1.2.5 → 1.2.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +179 -107
- package/README.pypi.md +179 -107
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
[](https://github.com/uno-km/termux-stt)
|
|
7
7
|
[](https://www.vulkan.org/)
|
|
8
8
|
|
|
9
|
-
> **Termux-STT** is an industrial-grade, zero-compilation on-device Speech-to-Text (STT) and multi-speaker diarization framework engineered specifically for Android Termux, ARM64 mobile hardware, and edge environments. By orchestrating a Tri-Engine acoustic pipeline (**Whisper.cpp**, **Vosk/Kaldi**, and **Sherpa-ONNX Zipformer**) with pure-Python x-vector speaker clustering and direct Vulkan GPU acceleration, Termux-STT achieves sub-realtime transcription speeds (up to **
|
|
9
|
+
> **Termux-STT** is an industrial-grade, zero-compilation on-device Speech-to-Text (STT) and multi-speaker diarization framework engineered specifically for Android Termux, ARM64 mobile hardware, and edge environments. By orchestrating a Tri-Engine acoustic pipeline (**Whisper.cpp**, **Vosk/Kaldi**, and **Sherpa-ONNX Zipformer**) with pure-Python x-vector speaker clustering and direct Vulkan GPU acceleration, Termux-STT achieves sub-realtime transcription speeds (up to **15.7x faster than real-time**, RTF **0.064x**) and continuous offline listening with zero cloud telemetry.
|
|
10
10
|
|
|
11
11
|
---
|
|
12
12
|
|
|
@@ -15,7 +15,7 @@
|
|
|
15
15
|
Termux-STT is distributed across both Python (PyPI) and Node.js (npm) ecosystems. It runs in unprivileged user-space on Android Termux (ARM64) and Linux aarch64/x86_64.
|
|
16
16
|
|
|
17
17
|
### 1.1 Prerequisites on Android Termux
|
|
18
|
-
Update package repositories and install
|
|
18
|
+
Update package repositories and install foundational audio and build utilities:
|
|
19
19
|
```bash
|
|
20
20
|
pkg update -y
|
|
21
21
|
pkg install -y clang python python-numpy nodejs termux-api ffmpeg pulseaudio
|
|
@@ -59,9 +59,9 @@ The automated engine installer (`EngineInstaller`) executes a 3-stage provisioni
|
|
|
59
59
|
1. **Native System Dependencies (`install_system_dependencies`)**:
|
|
60
60
|
Automatically invokes Termux `pkg` to install `ffmpeg`, `libbluray`, `libxml2`, `git`, `termux-api`, and `curl`, enabling universal audio decoding and microphone capture via Android APIs.
|
|
61
61
|
2. **Adaptive Engine Binary Provisioning (`install_whisper_cpp`)**:
|
|
62
|
-
-
|
|
63
|
-
-
|
|
64
|
-
-
|
|
62
|
+
- **Fast-Track Stream Extractor (~3s)**: On Android ARM64 Termux, precompiled Vulkan+NEON Bionic binaries (`whisper-cli-android-arm64.tar.gz`) are automatically extracted from GitHub Releases in ~3 seconds, completely eliminating 20-minute on-device compilation and mobile OOM aborts.
|
|
63
|
+
- **Zero-Hardcoding SSOT Endpoints**: Binary downloads dynamically route through unified SSOT candidate endpoints (`TERMUX_STT_RELEASE_TAG` -> `v{__version__}` -> `releases/latest/download` -> `uno-km/ameva-runtime` releases fallback) with automated fallback to on-device C++ compilation (`cmake` + `clang`) in offline or air-gapped environments.
|
|
64
|
+
- **Bundled Package Binary Fallback**: Automatically discovers and links bundled `termux_stt/bin/whisper-cli` if present.
|
|
65
65
|
3. **Sub-Engine Ecosystem Provisioning (`install_vosk`, `install_sherpa_onnx`)**:
|
|
66
66
|
Provisions `vosk` for sub-30ms real-time streaming and `sherpa-onnx` for next-generation ONNX Zipformer models, while pre-initializing model cache structures in `~/.cache/termux-stt/models/`.
|
|
67
67
|
|
|
@@ -95,7 +95,7 @@ pip install termux-stt ameva-runtime
|
|
|
95
95
|
npm install -g termux-stt @ameva/runtime
|
|
96
96
|
```
|
|
97
97
|
|
|
98
|
-
### 2.2 Hardware Diagnostics & Zero-Silent-Fallback
|
|
98
|
+
### 2.2 Hardware Diagnostics & Zero-Silent-Fallback Protocol
|
|
99
99
|
Verify Vulkan driver detection and SIMD feature availability:
|
|
100
100
|
```bash
|
|
101
101
|
termux-stt doctor
|
|
@@ -107,61 +107,106 @@ Termux-STT strictly enforces a **Zero-Silent-Fallback Protocol**:
|
|
|
107
107
|
|
|
108
108
|
---
|
|
109
109
|
|
|
110
|
-
## 3.
|
|
110
|
+
## 3. Comprehensive Usage Guide & Manual
|
|
111
111
|
|
|
112
|
-
Termux-STT provides intuitive interfaces across CLI, Python, and Node.js.
|
|
112
|
+
Termux-STT provides intuitive, high-performance interfaces across CLI, Python, and Node.js.
|
|
113
113
|
|
|
114
114
|
### 3.1 Command-Line Interface (CLI)
|
|
115
115
|
|
|
116
116
|
```bash
|
|
117
|
-
# 1.
|
|
117
|
+
# 1. Standard High-Accuracy Transcription (Greedy Search default: -bs 1)
|
|
118
118
|
termux-stt transcribe meeting.wav -e whisper -m base -l en
|
|
119
119
|
|
|
120
|
-
# 2.
|
|
120
|
+
# 2. Hardware-Accelerated Vulkan GPU Execution
|
|
121
|
+
termux-stt transcribe speech.wav -e whisper -m small -d vulkan
|
|
122
|
+
|
|
123
|
+
# 3. Multi-Beam Exploration for Precision Workloads (Explicit Beam Size Override)
|
|
124
|
+
termux-stt transcribe legal_deposition.wav -m small -bs 5
|
|
125
|
+
|
|
126
|
+
# 4. Voice Activity Detection (VAD) Pre-filtering (Drops Silence)
|
|
127
|
+
termux-stt transcribe lecture.wav -m base --vad
|
|
128
|
+
|
|
129
|
+
# 5. Export to Timestamped Subtitles (SRT / VTT / JSON)
|
|
121
130
|
termux-stt transcribe interview.mp3 --format srt -o output.srt
|
|
122
131
|
|
|
123
|
-
#
|
|
124
|
-
termux-stt
|
|
132
|
+
# 6. Audio Translation to English on the Fly
|
|
133
|
+
termux-stt transcribe interview_korean.wav -m small --translate -l ko
|
|
134
|
+
|
|
135
|
+
# 7. Multi-Speaker Diarization (Who Spoke When)
|
|
136
|
+
termux-stt diarize discussion.wav --speakers 3 --format rttm -o speakers.rttm
|
|
125
137
|
|
|
126
|
-
#
|
|
138
|
+
# 8. Live Microphone Real-Time Listening (Termux-API / Vosk)
|
|
127
139
|
termux-stt listen -e vosk -m small-ko
|
|
128
140
|
|
|
129
|
-
#
|
|
141
|
+
# 9. Zero-Configuration Built-in Benchmark Demo
|
|
130
142
|
termux-stt demo
|
|
143
|
+
|
|
144
|
+
# 10. Hardware & Driver Diagnostic Doctor
|
|
145
|
+
termux-stt doctor
|
|
146
|
+
|
|
147
|
+
# 11. Model Weight Management
|
|
148
|
+
termux-stt models list
|
|
149
|
+
termux-stt models download whisper small
|
|
131
150
|
```
|
|
132
151
|
|
|
133
152
|
### 3.2 Python SDK
|
|
153
|
+
|
|
134
154
|
```python
|
|
135
155
|
import termux_stt
|
|
136
156
|
|
|
137
|
-
# 1. Initialize High-Accuracy Whisper Engine with
|
|
138
|
-
engine = termux_stt.create_engine(
|
|
157
|
+
# 1. Initialize High-Accuracy Whisper Engine with Vulkan GPU Acceleration
|
|
158
|
+
engine = termux_stt.create_engine(
|
|
159
|
+
"whisper",
|
|
160
|
+
model="base",
|
|
161
|
+
device="auto", # "auto", "vulkan", "gpu", or "cpu"
|
|
162
|
+
beam_size=1 # Default: 1 (fast greedy decoding); set 5+ for multi-beam
|
|
163
|
+
)
|
|
139
164
|
|
|
140
|
-
# 2. Transcribe Audio File
|
|
141
|
-
result = engine.transcribe("meeting.wav")
|
|
165
|
+
# 2. Transcribe Audio File with Detailed Segment Output
|
|
166
|
+
result = engine.transcribe("meeting.wav", lang="en")
|
|
142
167
|
print(f"Full Text: {result.text}")
|
|
168
|
+
print(f"Duration: {result.audio_duration_sec:.2f}s | Elapsed: {result.elapsed_ms:.1f}ms | RTF: {result.rtf:.4f}x")
|
|
169
|
+
|
|
143
170
|
for segment in result.segments:
|
|
144
171
|
print(f"[{segment.start_sec:.2f}s -> {segment.end_sec:.2f}s] {segment.text}")
|
|
145
172
|
|
|
146
|
-
# 3.
|
|
173
|
+
# 3. Voice Activity Detection & Context Prompting
|
|
174
|
+
result_vad = engine.transcribe("noisy_lecture.wav", vad=True, prompt="Discussion on quantum computing")
|
|
175
|
+
print(f"VAD Result: {result_vad.text}")
|
|
176
|
+
|
|
177
|
+
# 4. Instant Low-Latency Streaming with Vosk Engine (<30ms Latency)
|
|
147
178
|
vosk_engine = termux_stt.create_engine("vosk", model="small-ko")
|
|
148
179
|
vosk_result = vosk_engine.transcribe("quick_voice.wav")
|
|
149
180
|
print(f"Vosk Output: {vosk_result.text}")
|
|
150
181
|
```
|
|
151
182
|
|
|
152
183
|
### 3.3 Node.js / TypeScript SDK
|
|
184
|
+
|
|
153
185
|
```typescript
|
|
154
186
|
import { createEngine } from 'termux-stt';
|
|
155
187
|
|
|
156
188
|
async function main() {
|
|
189
|
+
// Initialize Whisper engine with hardware acceleration
|
|
157
190
|
const engine = createEngine('whisper', {
|
|
158
191
|
model: 'base',
|
|
159
|
-
device: 'auto'
|
|
192
|
+
device: 'auto',
|
|
193
|
+
beamSize: 1
|
|
194
|
+
});
|
|
195
|
+
|
|
196
|
+
// Transcribe audio file
|
|
197
|
+
const result = await engine.transcribe('meeting.wav', {
|
|
198
|
+
lang: 'en',
|
|
199
|
+
format: 'json'
|
|
160
200
|
});
|
|
161
201
|
|
|
162
|
-
const result = await engine.transcribe('sample.wav');
|
|
163
202
|
console.log('Transcription:', result.text);
|
|
164
203
|
console.log(`Elapsed Time: ${result.elapsedMs}ms | RTF: ${result.rtf}x`);
|
|
204
|
+
|
|
205
|
+
// Stream partial results from microphone
|
|
206
|
+
const voskEngine = createEngine('vosk', { model: 'small-ko' });
|
|
207
|
+
voskEngine.on('transcript', (data) => {
|
|
208
|
+
console.log('Live Stream:', data.text);
|
|
209
|
+
});
|
|
165
210
|
}
|
|
166
211
|
|
|
167
212
|
main().catch(console.error);
|
|
@@ -169,7 +214,7 @@ main().catch(console.error);
|
|
|
169
214
|
|
|
170
215
|
---
|
|
171
216
|
|
|
172
|
-
## 4. Advanced
|
|
217
|
+
## 4. Advanced Architecture & Deep-Dive
|
|
173
218
|
|
|
174
219
|
Termux-STT features a versatile Tri-Engine architecture designed to adapt dynamically between studio precision and low-latency continuous listening.
|
|
175
220
|
|
|
@@ -191,7 +236,7 @@ flowchart TD
|
|
|
191
236
|
```
|
|
192
237
|
|
|
193
238
|
### 4.1 Subprocess Process Isolation & Mobile Crash Protection
|
|
194
|
-
Android Termux environments are prone to out-of-memory kernel kills (OOM) and SIGSEGV segmentation faults during heavy native C++ tensor inference. Termux-STT wraps `whisper.cpp` and `sherpa-onnx` in isolated process pools (`ProcessPool`), intercepting crashes gracefully without aborting the host Python application.
|
|
239
|
+
Android Termux environments are prone to out-of-memory kernel kills (OOM) and SIGSEGV segmentation faults during heavy native C++ tensor inference. Termux-STT wraps `whisper.cpp` and `sherpa-onnx` in isolated process pools (`ProcessPool`), intercepting crashes gracefully and returning typed exceptions (`AMEVA-STT-E002`) without aborting the host Python application.
|
|
195
240
|
|
|
196
241
|
### 4.2 Multi-Speaker Diarization Pipeline (No Scikit-Learn Needed)
|
|
197
242
|
Traditional speaker diarization requires heavy machine learning frameworks (`scikit-learn`, `torchaudio`). Termux-STT integrates an ultra-lightweight **HybridEngine**:
|
|
@@ -209,48 +254,48 @@ for seg in result.segments:
|
|
|
209
254
|
print(f"[{seg.speaker_id}] {seg.start_sec:.1f}s - {seg.end_sec:.1f}s: {seg.text}")
|
|
210
255
|
```
|
|
211
256
|
|
|
212
|
-
### 4.3
|
|
213
|
-
|
|
214
|
-
```python
|
|
215
|
-
import termux_stt
|
|
216
|
-
|
|
217
|
-
def on_partial_speech(text):
|
|
218
|
-
print(f"Live Stream: {text}", end="\r", flush=True)
|
|
219
|
-
|
|
220
|
-
engine = termux_stt.create_engine("vosk", model="small-ko")
|
|
221
|
-
# Listen continuously with Voice Activity Detection
|
|
222
|
-
engine.listen(callback=on_partial_speech, sample_rate=16000)
|
|
223
|
-
```
|
|
257
|
+
### 4.3 Mobile Production Greedy Search (`-bs 1`) Policy
|
|
258
|
+
Autoregressive decoding in Whisper generates text token-by-token. Standard desktop Whisper defaults to Beam Search ($B=5$), which maintains 5 parallel hypothesis states in memory. On mobile Unified Memory Architectures (UMA), this causes severe memory bus saturation, inflating the Key-Value (KV) cache from $49.8 ext{ MB}$ to $249 ext{ MB}$. Termux-STT hardcodes **Greedy Search (`--beam-size 1` / `-bs 1`)** as the mobile production default, slashing decoder memory bus traffic by **5x** while preserving full user sovereignty to override with `--beam-size 5`.
|
|
224
259
|
|
|
225
260
|
---
|
|
226
261
|
|
|
227
|
-
## 5. Feature & Parameter Matrix
|
|
262
|
+
## 5. Master Feature & Parameter Matrix
|
|
228
263
|
|
|
229
264
|
### 5.1 CLI Subcommands Overview
|
|
230
265
|
|
|
231
266
|
| Subcommand | Description | Example |
|
|
232
267
|
| :--- | :--- | :--- |
|
|
233
|
-
| `transcribe` | Transcribes audio file with chosen engine and
|
|
234
|
-
| `listen` | Captures live microphone audio and streams transcriptions. | `termux-stt listen -e vosk -m small-ko` |
|
|
235
|
-
| `diarize` | Identifies distinct speakers and outputs timestamped RTTM. | `termux-stt diarize meeting.wav --speakers 3` |
|
|
268
|
+
| `transcribe` | Transcribes audio file with chosen engine, model, and hardware backend. | `termux-stt transcribe speech.wav -e whisper -m base -d vulkan` |
|
|
269
|
+
| `listen` | Captures live microphone audio and streams real-time transcriptions. | `termux-stt listen -e vosk -m small-ko` |
|
|
270
|
+
| `diarize` | Identifies distinct speakers and outputs timestamped RTTM segmentation. | `termux-stt diarize meeting.wav --speakers 3` |
|
|
236
271
|
| `demo` | Runs end-to-end self-test on bundled JFK sample audio. | `termux-stt demo` |
|
|
237
|
-
| `doctor` | Diagnoses hardware SIMD, Vulkan GPU, and audio
|
|
238
|
-
| `benchmark` | Profiles Real-Time Factor (RTF) and memory allocation. | `termux-stt benchmark speech.wav` |
|
|
239
|
-
| `models` | Lists, downloads, and inspects cached offline weights. | `termux-stt models list` |
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
|
245
|
-
|
|
|
246
|
-
|
|
|
247
|
-
|
|
|
248
|
-
| `-d
|
|
249
|
-
|
|
|
250
|
-
| `--
|
|
251
|
-
| `--
|
|
252
|
-
| `--
|
|
253
|
-
|
|
|
272
|
+
| `doctor` | Diagnoses hardware SIMD, Vulkan GPU drivers, and audio subsystems. | `termux-stt doctor` |
|
|
273
|
+
| `benchmark` | Profiles Real-Time Factor (RTF), latency, and memory allocation. | `termux-stt benchmark --audio speech.wav` |
|
|
274
|
+
| `models` | Lists, downloads, and inspects cached offline model weights. | `termux-stt models list` |
|
|
275
|
+
| `install` | 1-Click automated installer for native binaries, codecs, and models. | `termux-stt install` |
|
|
276
|
+
|
|
277
|
+
### 5.2 Comprehensive Transcription Parameters Matrix
|
|
278
|
+
|
|
279
|
+
| Parameter Flag | Short | Type | Default | Description |
|
|
280
|
+
| :--- | :---: | :---: | :---: | :--- |
|
|
281
|
+
| `--engine` | `-e` | `enum` | `whisper` | Acoustic engine selection: `whisper`, `vosk`, `sherpa`, `hybrid`. |
|
|
282
|
+
| `--model` | `-m` | `string` | `base` | Model profile: `tiny`, `base`, `small`, `medium`, `turbo`, `small-ko`. |
|
|
283
|
+
| `--device` | `-d` / `-b` | `enum` | `auto` | Acceleration backend: `auto`, `vulkan`, `gpu`, `cpu`. |
|
|
284
|
+
| `--beam-size` | `-bs` | `int` | `1` | Beam search width. `1` = Greedy decoding (fastest mobile default); `5`+ = multi-beam. |
|
|
285
|
+
| `--lang` | `-l` | `string` | `ko` | Target language locale code (`en`, `ko`, `ja`, `zh`, `auto`). |
|
|
286
|
+
| `--vad` | | `flag` | `False` | Enables Voice Activity Detection (VAD) pre-filtering to skip silent intervals. |
|
|
287
|
+
| `--threads` | `-t` | `int` | *(Optimal)* | Number of worker threads pinned to ARM Cortex-X / big cores. |
|
|
288
|
+
| `--quantization` | | `enum` | `q5_1` | Model quantization level: `none`, `q4_0`, `q5_1`, `q8_0`, `f16`. |
|
|
289
|
+
| `--prompt` | | `string` | `None` | Initial prompt / context prefix passed to the decoder. |
|
|
290
|
+
| `--temperature` | | `float` | `0.0` | Sampling temperature for token decoding (lower is more deterministic). |
|
|
291
|
+
| `--translate` | | `flag` | `False` | Translates spoken audio directly into English text. |
|
|
292
|
+
| `--format` | | `enum` | `text` | Output formatting: `text`, `json`, `srt`, `vtt`, `rttm`. |
|
|
293
|
+
| `--output` | `-o` | `path` | `stdout` | Destination file path for generated transcript. |
|
|
294
|
+
| `--diarize` | | `flag` | `False` | Enables speaker identity clustering and segment alignment. |
|
|
295
|
+
| `--speakers` | | `int` | `2` | Expected number of speaker clusters for diarization. |
|
|
296
|
+
| `--demo` | | `flag` | `False` | Automatically uses bundled 60.00s JFK Inaugural Address benchmark audio. |
|
|
297
|
+
| `--extra-args` | | `string` | `None` | Raw CLI arguments passed directly to the underlying engine executable. |
|
|
298
|
+
| `--verbose` | | `flag` | `False` | Enables detailed debug and hardware telemetry logging. |
|
|
254
299
|
|
|
255
300
|
---
|
|
256
301
|
|
|
@@ -266,7 +311,7 @@ import termux_tts as tts
|
|
|
266
311
|
|
|
267
312
|
def run_conversational_cycle(user_audio="input.wav"):
|
|
268
313
|
# 1. Listen & Transcribe User Voice via Termux-STT
|
|
269
|
-
stt = termux_stt.create_engine("whisper", model="base")
|
|
314
|
+
stt = termux_stt.create_engine("whisper", model="base", device="auto", beam_size=1)
|
|
270
315
|
user_text = stt.transcribe(user_audio).text
|
|
271
316
|
print(f"Heard: {user_text}")
|
|
272
317
|
|
|
@@ -296,38 +341,63 @@ print(f"Installed Engines: {report.get('available_engines')}")
|
|
|
296
341
|
|
|
297
342
|
---
|
|
298
343
|
|
|
299
|
-
## 7.
|
|
300
|
-
|
|
301
|
-
### 7.1
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
| Target Device |
|
|
305
|
-
| :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
|
|
306
|
-
| **Galaxy S25** | Snapdragon 8 Elite
|
|
307
|
-
|
|
|
308
|
-
|
|
|
309
|
-
|
|
|
310
|
-
| **Galaxy
|
|
311
|
-
|
|
|
312
|
-
|
|
|
344
|
+
## 7. Empirical Mobile Hardware Benchmarks (6-SoC Physical Fleet)
|
|
345
|
+
|
|
346
|
+
### 7.1 Production Fleet Benchmark Scorecard
|
|
347
|
+
Empirical benchmarks conducted on physical Android hardware using JFK's 60.00s 16kHz Mono Inaugural Address (`samples/jfk_1min.wav`) with Greedy Search (`-bs 1`) under Vulkan GPU acceleration:
|
|
348
|
+
|
|
349
|
+
| Target Device | Silicon SoC | GPU Architecture | Model | Parameters | Processing Latency | RTF | Realtime Speed | Status |
|
|
350
|
+
| :--- | :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
|
|
351
|
+
| **Galaxy S25** | Snapdragon 8 Elite | Adreno 830 | Tiny | 39M | **3.82 s** | **0.064x** | **15.7x Realtime** | **PASS** |
|
|
352
|
+
| | | | Base | 74M | **7.14 s** | **0.119x** | **8.4x Realtime** | **PASS** |
|
|
353
|
+
| | | | Small | 244M | **26.63 s** | **0.444x** | **2.3x Realtime** | **PASS** |
|
|
354
|
+
| | | | Turbo | 809M | **68.42 s** | **1.140x** | **0.88x Realtime** | **PASS** |
|
|
355
|
+
| **Galaxy S22** | Snapdragon 8 Gen 1 | Adreno 730 | Tiny | 39M | **27.12 s** | **0.452x** | **2.21x Realtime** | **PASS** |
|
|
356
|
+
| | | | Base | 74M | **28.37 s** | **0.473x** | **2.11x Realtime** | **PASS** |
|
|
357
|
+
| | | | Small | 244M | **70.92 s** | **1.182x** | **0.85x Realtime** | **PASS** |
|
|
358
|
+
| | | | Turbo | 809M | **291.27 s** | **4.855x** | **0.21x Realtime** | **PASS** |
|
|
359
|
+
| **Galaxy S21** | Exynos 2100 | Mali-G78 MP14 | Tiny | 39M | **26.38 s** | **0.440x** | **2.27x Realtime** | **PASS** (Fleet Best Tiny) |
|
|
360
|
+
| | | | Base | 74M | **41.88 s** | **0.698x** | **1.43x Realtime** | **PASS** |
|
|
361
|
+
| | | | Small | 244M | **89.15 s** | **1.486x** | **0.67x Realtime** | **PASS** |
|
|
362
|
+
| | | | Turbo | 809M | **723.83 s** | **12.064x**| **0.08x Realtime** | **PASS** |
|
|
363
|
+
| **Galaxy A35** | Exynos 1380 | Mali-G68 MP5 | Tiny | 39M | **42.26 s** | **0.704x** | **1.42x Realtime** | **PASS** |
|
|
364
|
+
| | | | Base | 74M | **34.22 s** | **0.570x** | **1.75x Realtime** | **PASS** |
|
|
365
|
+
| | | | Small | 244M | **63.80 s** | **1.063x** | **0.94x Realtime** | **PASS** (Fleet Best Small) |
|
|
366
|
+
| | | | Turbo | 809M | **216.65 s** | **3.611x** | **0.28x Realtime** | **PASS** (Fleet Best Turbo) |
|
|
367
|
+
| **Galaxy A53** | Exynos 1280 | Mali-G68 MP4 | Tiny | 39M | **197.90 s** | **3.298x** | **0.30x Realtime** | **PASS** |
|
|
368
|
+
| | | | Base | 74M | **374.81 s** | **6.247x** | **0.16x Realtime** | **PASS** |
|
|
369
|
+
| | | | Small | 244M | **155.83 s** | **2.597x** | **0.39x Realtime** | **PASS** (+2GB RAM Plus) |
|
|
370
|
+
| | | | Turbo | 809M | **759.52 s** | **12.659x**| **0.08x Realtime** | **PASS** (+2GB RAM Plus) |
|
|
371
|
+
| **Galaxy S20** | Snapdragon 865 | Adreno 650 | Tiny | 39M | **30.04 s** | **0.501x** | **2.00x Realtime** | **PASS** |
|
|
372
|
+
| | | | Base | 74M | **45.71 s** | **0.762x** | **1.31x Realtime** | **PASS** |
|
|
373
|
+
| | | | Small | 244M | **128.50 s** | **2.142x** | **0.47x Realtime** | **PASS** |
|
|
374
|
+
| | | | Turbo | 809M | **271.02 s** | N/A | N/A | **FAIL** (KGSL Watchdog at 270s; no artificial block) |
|
|
313
375
|
|
|
314
376
|
> **Real-Time Factor (RTF) Definition**: $\text{RTF} = \frac{\text{Processing Latency (Seconds)}}{\text{Audio Duration (Seconds)}}$.
|
|
315
|
-
> An RTF of `0.
|
|
377
|
+
> An RTF of `0.064x` means 60 seconds of recorded speech is transcribed into text in only **3.82 seconds**.
|
|
378
|
+
|
|
379
|
+
### 7.2 CPU vs. Vulkan GPU Speedup Comparison
|
|
380
|
+
Comparative evaluation against ARM NEON 4-thread CPU execution illustrates massive GPU acceleration:
|
|
316
381
|
|
|
317
|
-
### 7.2 Verified Transcription Output Sample
|
|
318
382
|
```text
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
383
|
+
[Throughput Comparison: 60s JFK Audio Processing Speed]
|
|
384
|
+
Model: Whisper Small (244M)
|
|
385
|
+
|
|
386
|
+
Galaxy A35 (Exynos 1380 / Mali-G68 MP5)
|
|
387
|
+
CPU (NEON 4-threads): |==================================================| 482.1s
|
|
388
|
+
GPU (Vulkan0): |======| 63.8s [7.56x Faster]
|
|
323
389
|
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
390
|
+
Galaxy S22 (Snapdragon 8 Gen 1 / Adreno 730)
|
|
391
|
+
CPU (NEON 4-threads): |========================================| 394.3s
|
|
392
|
+
GPU (Vulkan0): |=======| 70.9s [5.56x Faster]
|
|
327
393
|
|
|
328
|
-
|
|
394
|
+
Galaxy S20 (Snapdragon 865 / Adreno 650)
|
|
395
|
+
CPU (NEON 4-threads): |==================================================| 512.6s
|
|
396
|
+
GPU (Vulkan0): |============| 128.5s [3.99x Faster]
|
|
329
397
|
```
|
|
330
398
|
|
|
399
|
+
Vulkan native GPU execution achieves a **$4.0\times\text{ to }7.6\times$ throughput speedup** over optimized multi-threaded ARM NEON CPU computation while reducing overall battery drain via the *Race-to-Sleep* principle.
|
|
400
|
+
|
|
331
401
|
---
|
|
332
402
|
|
|
333
403
|
## 8. GPU Interconnect Architecture & Compatibility
|
|
@@ -337,65 +407,65 @@ Termux-STT interfaces directly with Android's Bionic Vulkan loader (`/system/lib
|
|
|
337
407
|
|
|
338
408
|
### 8.2 Silicon Compatibility Matrix
|
|
339
409
|
- **Qualcomm Snapdragon (Adreno 6xx, 7xx, 8xx)**:
|
|
340
|
-
- **Tier-1 Full Support**. Native FP16 compute instructions and high dispatch concurrency deliver RTF performance as fast as **0.
|
|
410
|
+
- **Tier-1 Full Support**. Native FP16 compute instructions and high dispatch concurrency deliver RTF performance as fast as **0.064x** on Snapdragon 8 Elite. Adreno SoftMax workgroups are strictly constrained to hardware subgroup size (64) for rock-solid stability.
|
|
341
411
|
- **Samsung Exynos / MediaTek Dimensity (ARM Mali / Immortalis)**:
|
|
342
|
-
- **Supported**. Mali tile-based architectures benefit from
|
|
412
|
+
- **Supported**. Mali tile-based architectures benefit from continuous compute queues. Full execution across Tiny, Base, Small, and Turbo without shader stalls.
|
|
343
413
|
- **Strict Zero-Silent-Fallback**:
|
|
344
|
-
- Requesting `--device vulkan` without valid Vulkan drivers immediately triggers `PlatformNotSupportedError`, preventing silent fallback to unoptimized CPU execution.
|
|
414
|
+
- Requesting `--device vulkan` without valid Vulkan drivers immediately triggers typed exceptions (`PlatformNotSupportedError`, error code `AMEVA-STT-E002`), preventing silent fallback to unoptimized CPU execution.
|
|
345
415
|
|
|
346
416
|
---
|
|
347
417
|
|
|
348
|
-
##
|
|
418
|
+
## 9. Architectural Trade-Off Analysis & Decision Outcomes
|
|
419
|
+
|
|
420
|
+
In resource-constrained mobile systems engineering, every architectural design represents an explicit compromise between competing physical constraints:
|
|
349
421
|
|
|
350
|
-
|
|
|
351
|
-
| :--- | :--- | :--- | :--- |
|
|
352
|
-
| **
|
|
353
|
-
| **
|
|
354
|
-
| **
|
|
355
|
-
| **
|
|
356
|
-
| **
|
|
422
|
+
| Trade-Off Domain | Physical Constraints & Situation | Architectural Options Evaluated | Decision Executed | Empirical Outcome & Ground Truth |
|
|
423
|
+
| :--- | :--- | :--- | :--- | :--- |
|
|
424
|
+
| **1. Decoding Search Policy** | Autoregressive decoding saturates UMA memory bandwidth and inflates KV cache memory footprint. | **Option A**: Standard Beam Search ($B=5$) prioritizing hypothesis exploration.<br>**Option B**: Single-path Greedy Decoding ($B=1$). | **Selected Option B (Greedy $B=1$)** as production default; preserved Option A as explicit user override flag. | **5x reduction** in decoder memory bus traffic; KV cache shrunk from $249 ext{ MB}$ to $49.8 ext{ MB}$; zero measurable WER degradation on standard benchmark audio; eliminated mobile thermal spikes. |
|
|
425
|
+
| **2. Command Buffer Granularity** | 32-layer Turbo encoder takes $135.5 ext{s}$ per 30s window on Adreno 650, breaching the 270s KGSL TDR on 60s audio. | **Option A**: Fragment graph into 32-node sub-batches (`GGML_VK_MAX_NODES_PER_SUBMIT=32`).<br>**Option B**: Retain monolithic submission and enforce Fail-Fast. | **Selected Option B (Monolithic)** for production default; rejected forced fragmentation. | Subdividing submissions overloaded Qualcomm's 2021 kernel ringbuffer, triggering `IOCTL_KGSL_GPU_COMMAND: errno 35 (EDEADLK)` at $136.7 ext{s}$. Monolithic execution allows S22/S25/A35/A53 to achieve peak throughput without synchronization pipeline stalls. |
|
|
426
|
+
| **3. Heterogeneous Execution Routing** | High encoder latency suggested splitting encoder to GPU and decoder to CPU. | **Option A**: Dynamic asymmetric hybrid offloading across CPU/GPU.<br>**Option B**: Pure native Vulkan ABI pipeline binding. | **Selected Option B (Pure Native ABI)**; permanently blacklisted asymmetric offloading. | Avoided $7.68 ext{ MB}$ per layer inter-device tensor ping-pong over non-coherent UMA caches; eliminated thread synchronization jitter and delivered a unified, maintainable codebase. |
|
|
427
|
+
| **4. Hardware Support Boundaries** | Snapdragon 865 cannot complete 60s monolithic Turbo without TDR due to 2021 driver limitations. | **Option A**: Hardcode software check blocking Turbo model selection on S20.<br>**Option B**: Zero artificial restrictions, allow users complete freedom, let kernel return native exit codes. | **Selected Option B (Zero Artificial Blocks)** with full documentation transparency. | Preserves OpenSSF open-source compliance; allows users running shorter audio clips ($<30 ext{s}$) or custom kernels to execute Turbo unimpeded; maintains complete engineering honesty. |
|
|
428
|
+
| **5. Virtual Memory on 6GB SoCs** | Exynos 1280 (A53) has 6GB physical RAM; loading 809M Turbo caused severe swap thrashing ($>900 ext{s}$ timeout). | **Option A**: Restrict A53 to Small models only.<br>**Option B**: Provision +2GB zRAM swap backing store (RAM Plus) accepting minor compression overhead. | **Selected Option B (+2GB RAM Plus)** in operating system settings. | Reduced page fault rate $P_{fault} < 0.0001$; eliminated `kswapd0` CPU spin loops; Small finished in **$155.83 ext{s}$** and Turbo finished in **$759.52 ext{s}$** with $100\%$ transcript fidelity. |
|
|
357
429
|
|
|
358
|
-
|
|
430
|
+
> For the comprehensive academic analysis, kernel watchdog expiry mathematical proofs, and driver boundary traces, refer to the [Master Technical Treatise (English)](docs/research/on_device_vulkan_stt_master_treatise.md) and [연구 백서 (Korean)](docs/research/on_device_vulkan_stt_master_treatise_kor.md).
|
|
359
431
|
|
|
360
432
|
---
|
|
361
433
|
|
|
362
|
-
##
|
|
434
|
+
## 10. Hardware Requirements & Operational Limits
|
|
363
435
|
|
|
364
|
-
###
|
|
436
|
+
### 10.1 Hardware Specifications
|
|
365
437
|
|
|
366
438
|
| Specification Metric | Minimum Requirements | Recommended Production Spec |
|
|
367
439
|
| :--- | :--- | :--- |
|
|
368
440
|
| **Operating System** | Android 9.0+ (API level 28+) / Linux 5.4+ | Android 12.0+ (API level 31+) |
|
|
369
441
|
| **Architecture** | ARM64 (aarch64) or x86_64 | ARM64-v8a / v9a |
|
|
370
|
-
| **System RAM** | 2 GB Total Unified RAM | 4 GB+ Unified RAM |
|
|
442
|
+
| **System RAM** | 2 GB Total Unified RAM | 4 GB+ Unified RAM (6GB+ for Turbo) |
|
|
371
443
|
| **Storage Footprint** | 150 MB (Vosk) / 300 MB (Whisper Base) | 1 GB Free Flash Storage |
|
|
372
444
|
| **Audio Subsystem** | Termux-API Microphone Permissions | 16kHz PCM Audio Capture Support |
|
|
373
445
|
|
|
374
|
-
###
|
|
446
|
+
### 10.2 Operational Limits & Best Practices
|
|
375
447
|
- **32-Bit ARM (armeabi-v7a)**: Not supported for Whisper GPU neural inference. Use Vosk CFFI for legacy 32-bit hardware.
|
|
376
448
|
- **Microphone Permissions**: Real-time microphone listening (`termux-stt listen`) requires Android microphone permission granted to Termux: `termux-microphone-record`.
|
|
449
|
+
- **6GB RAM Mid-Range Devices**: For large models (Turbo 809M) on 6GB devices such as Galaxy A53, enable +2GB RAM Plus (zRAM) in Android settings to eliminate page thrashing.
|
|
377
450
|
|
|
378
451
|
---
|
|
379
452
|
|
|
380
|
-
##
|
|
453
|
+
## 11. 24/7 Unattended Background Execution Guide
|
|
381
454
|
|
|
382
|
-
Android aggressively kills background user-space processes inside Termux unless battery and process monitor policies are explicitly configured
|
|
455
|
+
Android aggressively kills background user-space processes inside Termux unless battery and process monitor policies are explicitly configured:
|
|
383
456
|
|
|
384
|
-
###
|
|
457
|
+
### 11.1 Stage 1: Termux Kernel Wake-Lock
|
|
385
458
|
Prevent the mobile CPU from entering low-power sleep states:
|
|
386
459
|
```bash
|
|
387
|
-
# Acquire persistent CPU wake-lock
|
|
388
460
|
termux-wake-lock
|
|
389
461
|
```
|
|
390
462
|
|
|
391
|
-
###
|
|
463
|
+
### 11.2 Stage 2: Android GUI Battery Optimization Exemption
|
|
392
464
|
1. Open **Android Settings > Apps > Termux > Battery**.
|
|
393
465
|
2. Set battery policy to **Unrestricted** (Disable power-saving restrictions).
|
|
394
466
|
3. Under **Permissions**, grant **Microphone** and **Display over other apps**.
|
|
395
467
|
|
|
396
|
-
###
|
|
397
|
-
Android 12+ terminates background processes exceeding child process thresholds. Execute these commands via ADB:
|
|
398
|
-
|
|
468
|
+
### 11.3 Stage 3: ADB Phantom Process Killer Exemption (Android 12+)
|
|
399
469
|
```bash
|
|
400
470
|
# Disable Android Phantom Process Killer
|
|
401
471
|
adb shell device_config put activity_manager max_phantom_processes 2147483647
|
|
@@ -408,7 +478,7 @@ adb shell settings get global settings_enable_monitor_phantom_procs
|
|
|
408
478
|
|
|
409
479
|
---
|
|
410
480
|
|
|
411
|
-
##
|
|
481
|
+
## 12. Open Source License
|
|
412
482
|
|
|
413
483
|
Termux-STT is open-sourced under the **MIT License**.
|
|
414
484
|
|
|
@@ -436,7 +506,7 @@ SOFTWARE.
|
|
|
436
506
|
|
|
437
507
|
---
|
|
438
508
|
|
|
439
|
-
##
|
|
509
|
+
## 13. SEO Technical Keywords & Ecosystem Metadata
|
|
440
510
|
|
|
441
511
|
`termux`, `stt`, `speech-to-text`, `whisper`, `whisper-cpp`, `vosk`, `sherpa-onnx`, `diarization`, `speaker-diarization`, `voice-recognition`, `audio-transcription`, `on-device-ai`, `edge-ai`, `mobile-ai`, `vulkan`, `vulkan-compute`, `gpu-acceleration`, `real-time-factor`, `low-latency`, `arm64`, `android`, `snapdragon`, `adreno`, `exynos`, `arm-mali`, `zero-compilation`, `vad`, `voice-activity-detection`, `silero-vad`, `x-vector`, `k-means`, `clustering`, `srt-export`, `vtt-export`, `rttm`, `microphone-streaming`, `headless-audio`, `pulseaudio`, `offline-speech`, `privacy-first`, `termux-aichain`, `termux-tts`, `termux-llamacpp`, `termux-diffusion`, `ameva-runtime`, `ggml`, `quantization`, `bionic-libc`, `autonomous-agents`, `voice-assistant`
|
|
442
512
|
|
|
@@ -446,3 +516,5 @@ SOFTWARE.
|
|
|
446
516
|
- **Official Documentation Portal**: [https://uno-km.github.io/termux-stt/](https://uno-km.github.io/termux-stt/)
|
|
447
517
|
- **GitHub Repository**: [https://github.com/uno-km/termux-stt](https://github.com/uno-km/termux-stt)
|
|
448
518
|
- **AMEVA Foundation Portal**: [https://uno-km.vercel.app/foundation/index.html](https://uno-km.vercel.app/foundation/index.html)
|
|
519
|
+
- **Master Research Treatise (English)**: [docs/research/on_device_vulkan_stt_master_treatise.md](docs/research/on_device_vulkan_stt_master_treatise.md)
|
|
520
|
+
- **연구 백서 (Korean)**: [docs/research/on_device_vulkan_stt_master_treatise_kor.md](docs/research/on_device_vulkan_stt_master_treatise_kor.md)
|
package/README.pypi.md
CHANGED
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
[](https://github.com/uno-km/termux-stt)
|
|
7
7
|
[](https://www.vulkan.org/)
|
|
8
8
|
|
|
9
|
-
> **Termux-STT** is an industrial-grade, zero-compilation on-device Speech-to-Text (STT) and multi-speaker diarization framework engineered specifically for Android Termux, ARM64 mobile hardware, and edge environments. By orchestrating a Tri-Engine acoustic pipeline (**Whisper.cpp**, **Vosk/Kaldi**, and **Sherpa-ONNX Zipformer**) with pure-Python x-vector speaker clustering and direct Vulkan GPU acceleration, Termux-STT achieves sub-realtime transcription speeds (up to **
|
|
9
|
+
> **Termux-STT** is an industrial-grade, zero-compilation on-device Speech-to-Text (STT) and multi-speaker diarization framework engineered specifically for Android Termux, ARM64 mobile hardware, and edge environments. By orchestrating a Tri-Engine acoustic pipeline (**Whisper.cpp**, **Vosk/Kaldi**, and **Sherpa-ONNX Zipformer**) with pure-Python x-vector speaker clustering and direct Vulkan GPU acceleration, Termux-STT achieves sub-realtime transcription speeds (up to **15.7x faster than real-time**, RTF **0.064x**) and continuous offline listening with zero cloud telemetry.
|
|
10
10
|
|
|
11
11
|
---
|
|
12
12
|
|
|
@@ -15,7 +15,7 @@
|
|
|
15
15
|
Termux-STT is distributed across both Python (PyPI) and Node.js (npm) ecosystems. It runs in unprivileged user-space on Android Termux (ARM64) and Linux aarch64/x86_64.
|
|
16
16
|
|
|
17
17
|
### 1.1 Prerequisites on Android Termux
|
|
18
|
-
Update package repositories and install
|
|
18
|
+
Update package repositories and install foundational audio and build utilities:
|
|
19
19
|
```bash
|
|
20
20
|
pkg update -y
|
|
21
21
|
pkg install -y clang python python-numpy nodejs termux-api ffmpeg pulseaudio
|
|
@@ -59,9 +59,9 @@ The automated engine installer (`EngineInstaller`) executes a 3-stage provisioni
|
|
|
59
59
|
1. **Native System Dependencies (`install_system_dependencies`)**:
|
|
60
60
|
Automatically invokes Termux `pkg` to install `ffmpeg`, `libbluray`, `libxml2`, `git`, `termux-api`, and `curl`, enabling universal audio decoding and microphone capture via Android APIs.
|
|
61
61
|
2. **Adaptive Engine Binary Provisioning (`install_whisper_cpp`)**:
|
|
62
|
-
-
|
|
63
|
-
-
|
|
64
|
-
-
|
|
62
|
+
- **Fast-Track Stream Extractor (~3s)**: On Android ARM64 Termux, precompiled Vulkan+NEON Bionic binaries (`whisper-cli-android-arm64.tar.gz`) are automatically extracted from GitHub Releases in ~3 seconds, completely eliminating 20-minute on-device compilation and mobile OOM aborts.
|
|
63
|
+
- **Zero-Hardcoding SSOT Endpoints**: Binary downloads dynamically route through unified SSOT candidate endpoints (`TERMUX_STT_RELEASE_TAG` -> `v{__version__}` -> `releases/latest/download` -> `uno-km/ameva-runtime` releases fallback) with automated fallback to on-device C++ compilation (`cmake` + `clang`) in offline or air-gapped environments.
|
|
64
|
+
- **Bundled Package Binary Fallback**: Automatically discovers and links bundled `termux_stt/bin/whisper-cli` if present.
|
|
65
65
|
3. **Sub-Engine Ecosystem Provisioning (`install_vosk`, `install_sherpa_onnx`)**:
|
|
66
66
|
Provisions `vosk` for sub-30ms real-time streaming and `sherpa-onnx` for next-generation ONNX Zipformer models, while pre-initializing model cache structures in `~/.cache/termux-stt/models/`.
|
|
67
67
|
|
|
@@ -95,7 +95,7 @@ pip install termux-stt ameva-runtime
|
|
|
95
95
|
npm install -g termux-stt @ameva/runtime
|
|
96
96
|
```
|
|
97
97
|
|
|
98
|
-
### 2.2 Hardware Diagnostics & Zero-Silent-Fallback
|
|
98
|
+
### 2.2 Hardware Diagnostics & Zero-Silent-Fallback Protocol
|
|
99
99
|
Verify Vulkan driver detection and SIMD feature availability:
|
|
100
100
|
```bash
|
|
101
101
|
termux-stt doctor
|
|
@@ -107,61 +107,106 @@ Termux-STT strictly enforces a **Zero-Silent-Fallback Protocol**:
|
|
|
107
107
|
|
|
108
108
|
---
|
|
109
109
|
|
|
110
|
-
## 3.
|
|
110
|
+
## 3. Comprehensive Usage Guide & Manual
|
|
111
111
|
|
|
112
|
-
Termux-STT provides intuitive interfaces across CLI, Python, and Node.js.
|
|
112
|
+
Termux-STT provides intuitive, high-performance interfaces across CLI, Python, and Node.js.
|
|
113
113
|
|
|
114
114
|
### 3.1 Command-Line Interface (CLI)
|
|
115
115
|
|
|
116
116
|
```bash
|
|
117
|
-
# 1.
|
|
117
|
+
# 1. Standard High-Accuracy Transcription (Greedy Search default: -bs 1)
|
|
118
118
|
termux-stt transcribe meeting.wav -e whisper -m base -l en
|
|
119
119
|
|
|
120
|
-
# 2.
|
|
120
|
+
# 2. Hardware-Accelerated Vulkan GPU Execution
|
|
121
|
+
termux-stt transcribe speech.wav -e whisper -m small -d vulkan
|
|
122
|
+
|
|
123
|
+
# 3. Multi-Beam Exploration for Precision Workloads (Explicit Beam Size Override)
|
|
124
|
+
termux-stt transcribe legal_deposition.wav -m small -bs 5
|
|
125
|
+
|
|
126
|
+
# 4. Voice Activity Detection (VAD) Pre-filtering (Drops Silence)
|
|
127
|
+
termux-stt transcribe lecture.wav -m base --vad
|
|
128
|
+
|
|
129
|
+
# 5. Export to Timestamped Subtitles (SRT / VTT / JSON)
|
|
121
130
|
termux-stt transcribe interview.mp3 --format srt -o output.srt
|
|
122
131
|
|
|
123
|
-
#
|
|
124
|
-
termux-stt
|
|
132
|
+
# 6. Audio Translation to English on the Fly
|
|
133
|
+
termux-stt transcribe interview_korean.wav -m small --translate -l ko
|
|
134
|
+
|
|
135
|
+
# 7. Multi-Speaker Diarization (Who Spoke When)
|
|
136
|
+
termux-stt diarize discussion.wav --speakers 3 --format rttm -o speakers.rttm
|
|
125
137
|
|
|
126
|
-
#
|
|
138
|
+
# 8. Live Microphone Real-Time Listening (Termux-API / Vosk)
|
|
127
139
|
termux-stt listen -e vosk -m small-ko
|
|
128
140
|
|
|
129
|
-
#
|
|
141
|
+
# 9. Zero-Configuration Built-in Benchmark Demo
|
|
130
142
|
termux-stt demo
|
|
143
|
+
|
|
144
|
+
# 10. Hardware & Driver Diagnostic Doctor
|
|
145
|
+
termux-stt doctor
|
|
146
|
+
|
|
147
|
+
# 11. Model Weight Management
|
|
148
|
+
termux-stt models list
|
|
149
|
+
termux-stt models download whisper small
|
|
131
150
|
```
|
|
132
151
|
|
|
133
152
|
### 3.2 Python SDK
|
|
153
|
+
|
|
134
154
|
```python
|
|
135
155
|
import termux_stt
|
|
136
156
|
|
|
137
|
-
# 1. Initialize High-Accuracy Whisper Engine with
|
|
138
|
-
engine = termux_stt.create_engine(
|
|
157
|
+
# 1. Initialize High-Accuracy Whisper Engine with Vulkan GPU Acceleration
|
|
158
|
+
engine = termux_stt.create_engine(
|
|
159
|
+
"whisper",
|
|
160
|
+
model="base",
|
|
161
|
+
device="auto", # "auto", "vulkan", "gpu", or "cpu"
|
|
162
|
+
beam_size=1 # Default: 1 (fast greedy decoding); set 5+ for multi-beam
|
|
163
|
+
)
|
|
139
164
|
|
|
140
|
-
# 2. Transcribe Audio File
|
|
141
|
-
result = engine.transcribe("meeting.wav")
|
|
165
|
+
# 2. Transcribe Audio File with Detailed Segment Output
|
|
166
|
+
result = engine.transcribe("meeting.wav", lang="en")
|
|
142
167
|
print(f"Full Text: {result.text}")
|
|
168
|
+
print(f"Duration: {result.audio_duration_sec:.2f}s | Elapsed: {result.elapsed_ms:.1f}ms | RTF: {result.rtf:.4f}x")
|
|
169
|
+
|
|
143
170
|
for segment in result.segments:
|
|
144
171
|
print(f"[{segment.start_sec:.2f}s -> {segment.end_sec:.2f}s] {segment.text}")
|
|
145
172
|
|
|
146
|
-
# 3.
|
|
173
|
+
# 3. Voice Activity Detection & Context Prompting
|
|
174
|
+
result_vad = engine.transcribe("noisy_lecture.wav", vad=True, prompt="Discussion on quantum computing")
|
|
175
|
+
print(f"VAD Result: {result_vad.text}")
|
|
176
|
+
|
|
177
|
+
# 4. Instant Low-Latency Streaming with Vosk Engine (<30ms Latency)
|
|
147
178
|
vosk_engine = termux_stt.create_engine("vosk", model="small-ko")
|
|
148
179
|
vosk_result = vosk_engine.transcribe("quick_voice.wav")
|
|
149
180
|
print(f"Vosk Output: {vosk_result.text}")
|
|
150
181
|
```
|
|
151
182
|
|
|
152
183
|
### 3.3 Node.js / TypeScript SDK
|
|
184
|
+
|
|
153
185
|
```typescript
|
|
154
186
|
import { createEngine } from 'termux-stt';
|
|
155
187
|
|
|
156
188
|
async function main() {
|
|
189
|
+
// Initialize Whisper engine with hardware acceleration
|
|
157
190
|
const engine = createEngine('whisper', {
|
|
158
191
|
model: 'base',
|
|
159
|
-
device: 'auto'
|
|
192
|
+
device: 'auto',
|
|
193
|
+
beamSize: 1
|
|
194
|
+
});
|
|
195
|
+
|
|
196
|
+
// Transcribe audio file
|
|
197
|
+
const result = await engine.transcribe('meeting.wav', {
|
|
198
|
+
lang: 'en',
|
|
199
|
+
format: 'json'
|
|
160
200
|
});
|
|
161
201
|
|
|
162
|
-
const result = await engine.transcribe('sample.wav');
|
|
163
202
|
console.log('Transcription:', result.text);
|
|
164
203
|
console.log(`Elapsed Time: ${result.elapsedMs}ms | RTF: ${result.rtf}x`);
|
|
204
|
+
|
|
205
|
+
// Stream partial results from microphone
|
|
206
|
+
const voskEngine = createEngine('vosk', { model: 'small-ko' });
|
|
207
|
+
voskEngine.on('transcript', (data) => {
|
|
208
|
+
console.log('Live Stream:', data.text);
|
|
209
|
+
});
|
|
165
210
|
}
|
|
166
211
|
|
|
167
212
|
main().catch(console.error);
|
|
@@ -169,7 +214,7 @@ main().catch(console.error);
|
|
|
169
214
|
|
|
170
215
|
---
|
|
171
216
|
|
|
172
|
-
## 4. Advanced
|
|
217
|
+
## 4. Advanced Architecture & Deep-Dive
|
|
173
218
|
|
|
174
219
|
Termux-STT features a versatile Tri-Engine architecture designed to adapt dynamically between studio precision and low-latency continuous listening.
|
|
175
220
|
|
|
@@ -191,7 +236,7 @@ flowchart TD
|
|
|
191
236
|
```
|
|
192
237
|
|
|
193
238
|
### 4.1 Subprocess Process Isolation & Mobile Crash Protection
|
|
194
|
-
Android Termux environments are prone to out-of-memory kernel kills (OOM) and SIGSEGV segmentation faults during heavy native C++ tensor inference. Termux-STT wraps `whisper.cpp` and `sherpa-onnx` in isolated process pools (`ProcessPool`), intercepting crashes gracefully without aborting the host Python application.
|
|
239
|
+
Android Termux environments are prone to out-of-memory kernel kills (OOM) and SIGSEGV segmentation faults during heavy native C++ tensor inference. Termux-STT wraps `whisper.cpp` and `sherpa-onnx` in isolated process pools (`ProcessPool`), intercepting crashes gracefully and returning typed exceptions (`AMEVA-STT-E002`) without aborting the host Python application.
|
|
195
240
|
|
|
196
241
|
### 4.2 Multi-Speaker Diarization Pipeline (No Scikit-Learn Needed)
|
|
197
242
|
Traditional speaker diarization requires heavy machine learning frameworks (`scikit-learn`, `torchaudio`). Termux-STT integrates an ultra-lightweight **HybridEngine**:
|
|
@@ -209,48 +254,48 @@ for seg in result.segments:
|
|
|
209
254
|
print(f"[{seg.speaker_id}] {seg.start_sec:.1f}s - {seg.end_sec:.1f}s: {seg.text}")
|
|
210
255
|
```
|
|
211
256
|
|
|
212
|
-
### 4.3
|
|
213
|
-
|
|
214
|
-
```python
|
|
215
|
-
import termux_stt
|
|
216
|
-
|
|
217
|
-
def on_partial_speech(text):
|
|
218
|
-
print(f"Live Stream: {text}", end="\r", flush=True)
|
|
219
|
-
|
|
220
|
-
engine = termux_stt.create_engine("vosk", model="small-ko")
|
|
221
|
-
# Listen continuously with Voice Activity Detection
|
|
222
|
-
engine.listen(callback=on_partial_speech, sample_rate=16000)
|
|
223
|
-
```
|
|
257
|
+
### 4.3 Mobile Production Greedy Search (`-bs 1`) Policy
|
|
258
|
+
Autoregressive decoding in Whisper generates text token-by-token. Standard desktop Whisper defaults to Beam Search ($B=5$), which maintains 5 parallel hypothesis states in memory. On mobile Unified Memory Architectures (UMA), this causes severe memory bus saturation, inflating the Key-Value (KV) cache from $49.8 ext{ MB}$ to $249 ext{ MB}$. Termux-STT hardcodes **Greedy Search (`--beam-size 1` / `-bs 1`)** as the mobile production default, slashing decoder memory bus traffic by **5x** while preserving full user sovereignty to override with `--beam-size 5`.
|
|
224
259
|
|
|
225
260
|
---
|
|
226
261
|
|
|
227
|
-
## 5. Feature & Parameter Matrix
|
|
262
|
+
## 5. Master Feature & Parameter Matrix
|
|
228
263
|
|
|
229
264
|
### 5.1 CLI Subcommands Overview
|
|
230
265
|
|
|
231
266
|
| Subcommand | Description | Example |
|
|
232
267
|
| :--- | :--- | :--- |
|
|
233
|
-
| `transcribe` | Transcribes audio file with chosen engine and
|
|
234
|
-
| `listen` | Captures live microphone audio and streams transcriptions. | `termux-stt listen -e vosk -m small-ko` |
|
|
235
|
-
| `diarize` | Identifies distinct speakers and outputs timestamped RTTM. | `termux-stt diarize meeting.wav --speakers 3` |
|
|
268
|
+
| `transcribe` | Transcribes audio file with chosen engine, model, and hardware backend. | `termux-stt transcribe speech.wav -e whisper -m base -d vulkan` |
|
|
269
|
+
| `listen` | Captures live microphone audio and streams real-time transcriptions. | `termux-stt listen -e vosk -m small-ko` |
|
|
270
|
+
| `diarize` | Identifies distinct speakers and outputs timestamped RTTM segmentation. | `termux-stt diarize meeting.wav --speakers 3` |
|
|
236
271
|
| `demo` | Runs end-to-end self-test on bundled JFK sample audio. | `termux-stt demo` |
|
|
237
|
-
| `doctor` | Diagnoses hardware SIMD, Vulkan GPU, and audio
|
|
238
|
-
| `benchmark` | Profiles Real-Time Factor (RTF) and memory allocation. | `termux-stt benchmark speech.wav` |
|
|
239
|
-
| `models` | Lists, downloads, and inspects cached offline weights. | `termux-stt models list` |
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
|
245
|
-
|
|
|
246
|
-
|
|
|
247
|
-
|
|
|
248
|
-
| `-d
|
|
249
|
-
|
|
|
250
|
-
| `--
|
|
251
|
-
| `--
|
|
252
|
-
| `--
|
|
253
|
-
|
|
|
272
|
+
| `doctor` | Diagnoses hardware SIMD, Vulkan GPU drivers, and audio subsystems. | `termux-stt doctor` |
|
|
273
|
+
| `benchmark` | Profiles Real-Time Factor (RTF), latency, and memory allocation. | `termux-stt benchmark --audio speech.wav` |
|
|
274
|
+
| `models` | Lists, downloads, and inspects cached offline model weights. | `termux-stt models list` |
|
|
275
|
+
| `install` | 1-Click automated installer for native binaries, codecs, and models. | `termux-stt install` |
|
|
276
|
+
|
|
277
|
+
### 5.2 Comprehensive Transcription Parameters Matrix
|
|
278
|
+
|
|
279
|
+
| Parameter Flag | Short | Type | Default | Description |
|
|
280
|
+
| :--- | :---: | :---: | :---: | :--- |
|
|
281
|
+
| `--engine` | `-e` | `enum` | `whisper` | Acoustic engine selection: `whisper`, `vosk`, `sherpa`, `hybrid`. |
|
|
282
|
+
| `--model` | `-m` | `string` | `base` | Model profile: `tiny`, `base`, `small`, `medium`, `turbo`, `small-ko`. |
|
|
283
|
+
| `--device` | `-d` / `-b` | `enum` | `auto` | Acceleration backend: `auto`, `vulkan`, `gpu`, `cpu`. |
|
|
284
|
+
| `--beam-size` | `-bs` | `int` | `1` | Beam search width. `1` = Greedy decoding (fastest mobile default); `5`+ = multi-beam. |
|
|
285
|
+
| `--lang` | `-l` | `string` | `ko` | Target language locale code (`en`, `ko`, `ja`, `zh`, `auto`). |
|
|
286
|
+
| `--vad` | | `flag` | `False` | Enables Voice Activity Detection (VAD) pre-filtering to skip silent intervals. |
|
|
287
|
+
| `--threads` | `-t` | `int` | *(Optimal)* | Number of worker threads pinned to ARM Cortex-X / big cores. |
|
|
288
|
+
| `--quantization` | | `enum` | `q5_1` | Model quantization level: `none`, `q4_0`, `q5_1`, `q8_0`, `f16`. |
|
|
289
|
+
| `--prompt` | | `string` | `None` | Initial prompt / context prefix passed to the decoder. |
|
|
290
|
+
| `--temperature` | | `float` | `0.0` | Sampling temperature for token decoding (lower is more deterministic). |
|
|
291
|
+
| `--translate` | | `flag` | `False` | Translates spoken audio directly into English text. |
|
|
292
|
+
| `--format` | | `enum` | `text` | Output formatting: `text`, `json`, `srt`, `vtt`, `rttm`. |
|
|
293
|
+
| `--output` | `-o` | `path` | `stdout` | Destination file path for generated transcript. |
|
|
294
|
+
| `--diarize` | | `flag` | `False` | Enables speaker identity clustering and segment alignment. |
|
|
295
|
+
| `--speakers` | | `int` | `2` | Expected number of speaker clusters for diarization. |
|
|
296
|
+
| `--demo` | | `flag` | `False` | Automatically uses bundled 60.00s JFK Inaugural Address benchmark audio. |
|
|
297
|
+
| `--extra-args` | | `string` | `None` | Raw CLI arguments passed directly to the underlying engine executable. |
|
|
298
|
+
| `--verbose` | | `flag` | `False` | Enables detailed debug and hardware telemetry logging. |
|
|
254
299
|
|
|
255
300
|
---
|
|
256
301
|
|
|
@@ -266,7 +311,7 @@ import termux_tts as tts
|
|
|
266
311
|
|
|
267
312
|
def run_conversational_cycle(user_audio="input.wav"):
|
|
268
313
|
# 1. Listen & Transcribe User Voice via Termux-STT
|
|
269
|
-
stt = termux_stt.create_engine("whisper", model="base")
|
|
314
|
+
stt = termux_stt.create_engine("whisper", model="base", device="auto", beam_size=1)
|
|
270
315
|
user_text = stt.transcribe(user_audio).text
|
|
271
316
|
print(f"Heard: {user_text}")
|
|
272
317
|
|
|
@@ -296,38 +341,63 @@ print(f"Installed Engines: {report.get('available_engines')}")
|
|
|
296
341
|
|
|
297
342
|
---
|
|
298
343
|
|
|
299
|
-
## 7.
|
|
300
|
-
|
|
301
|
-
### 7.1
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
| Target Device |
|
|
305
|
-
| :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
|
|
306
|
-
| **Galaxy S25** | Snapdragon 8 Elite
|
|
307
|
-
|
|
|
308
|
-
|
|
|
309
|
-
|
|
|
310
|
-
| **Galaxy
|
|
311
|
-
|
|
|
312
|
-
|
|
|
344
|
+
## 7. Empirical Mobile Hardware Benchmarks (6-SoC Physical Fleet)
|
|
345
|
+
|
|
346
|
+
### 7.1 Production Fleet Benchmark Scorecard
|
|
347
|
+
Empirical benchmarks conducted on physical Android hardware using JFK's 60.00s 16kHz Mono Inaugural Address (`samples/jfk_1min.wav`) with Greedy Search (`-bs 1`) under Vulkan GPU acceleration:
|
|
348
|
+
|
|
349
|
+
| Target Device | Silicon SoC | GPU Architecture | Model | Parameters | Processing Latency | RTF | Realtime Speed | Status |
|
|
350
|
+
| :--- | :--- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
|
|
351
|
+
| **Galaxy S25** | Snapdragon 8 Elite | Adreno 830 | Tiny | 39M | **3.82 s** | **0.064x** | **15.7x Realtime** | **PASS** |
|
|
352
|
+
| | | | Base | 74M | **7.14 s** | **0.119x** | **8.4x Realtime** | **PASS** |
|
|
353
|
+
| | | | Small | 244M | **26.63 s** | **0.444x** | **2.3x Realtime** | **PASS** |
|
|
354
|
+
| | | | Turbo | 809M | **68.42 s** | **1.140x** | **0.88x Realtime** | **PASS** |
|
|
355
|
+
| **Galaxy S22** | Snapdragon 8 Gen 1 | Adreno 730 | Tiny | 39M | **27.12 s** | **0.452x** | **2.21x Realtime** | **PASS** |
|
|
356
|
+
| | | | Base | 74M | **28.37 s** | **0.473x** | **2.11x Realtime** | **PASS** |
|
|
357
|
+
| | | | Small | 244M | **70.92 s** | **1.182x** | **0.85x Realtime** | **PASS** |
|
|
358
|
+
| | | | Turbo | 809M | **291.27 s** | **4.855x** | **0.21x Realtime** | **PASS** |
|
|
359
|
+
| **Galaxy S21** | Exynos 2100 | Mali-G78 MP14 | Tiny | 39M | **26.38 s** | **0.440x** | **2.27x Realtime** | **PASS** (Fleet Best Tiny) |
|
|
360
|
+
| | | | Base | 74M | **41.88 s** | **0.698x** | **1.43x Realtime** | **PASS** |
|
|
361
|
+
| | | | Small | 244M | **89.15 s** | **1.486x** | **0.67x Realtime** | **PASS** |
|
|
362
|
+
| | | | Turbo | 809M | **723.83 s** | **12.064x**| **0.08x Realtime** | **PASS** |
|
|
363
|
+
| **Galaxy A35** | Exynos 1380 | Mali-G68 MP5 | Tiny | 39M | **42.26 s** | **0.704x** | **1.42x Realtime** | **PASS** |
|
|
364
|
+
| | | | Base | 74M | **34.22 s** | **0.570x** | **1.75x Realtime** | **PASS** |
|
|
365
|
+
| | | | Small | 244M | **63.80 s** | **1.063x** | **0.94x Realtime** | **PASS** (Fleet Best Small) |
|
|
366
|
+
| | | | Turbo | 809M | **216.65 s** | **3.611x** | **0.28x Realtime** | **PASS** (Fleet Best Turbo) |
|
|
367
|
+
| **Galaxy A53** | Exynos 1280 | Mali-G68 MP4 | Tiny | 39M | **197.90 s** | **3.298x** | **0.30x Realtime** | **PASS** |
|
|
368
|
+
| | | | Base | 74M | **374.81 s** | **6.247x** | **0.16x Realtime** | **PASS** |
|
|
369
|
+
| | | | Small | 244M | **155.83 s** | **2.597x** | **0.39x Realtime** | **PASS** (+2GB RAM Plus) |
|
|
370
|
+
| | | | Turbo | 809M | **759.52 s** | **12.659x**| **0.08x Realtime** | **PASS** (+2GB RAM Plus) |
|
|
371
|
+
| **Galaxy S20** | Snapdragon 865 | Adreno 650 | Tiny | 39M | **30.04 s** | **0.501x** | **2.00x Realtime** | **PASS** |
|
|
372
|
+
| | | | Base | 74M | **45.71 s** | **0.762x** | **1.31x Realtime** | **PASS** |
|
|
373
|
+
| | | | Small | 244M | **128.50 s** | **2.142x** | **0.47x Realtime** | **PASS** |
|
|
374
|
+
| | | | Turbo | 809M | **271.02 s** | N/A | N/A | **FAIL** (KGSL Watchdog at 270s; no artificial block) |
|
|
313
375
|
|
|
314
376
|
> **Real-Time Factor (RTF) Definition**: $\text{RTF} = \frac{\text{Processing Latency (Seconds)}}{\text{Audio Duration (Seconds)}}$.
|
|
315
|
-
> An RTF of `0.
|
|
377
|
+
> An RTF of `0.064x` means 60 seconds of recorded speech is transcribed into text in only **3.82 seconds**.
|
|
378
|
+
|
|
379
|
+
### 7.2 CPU vs. Vulkan GPU Speedup Comparison
|
|
380
|
+
Comparative evaluation against ARM NEON 4-thread CPU execution illustrates massive GPU acceleration:
|
|
316
381
|
|
|
317
|
-
### 7.2 Verified Transcription Output Sample
|
|
318
382
|
```text
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
383
|
+
[Throughput Comparison: 60s JFK Audio Processing Speed]
|
|
384
|
+
Model: Whisper Small (244M)
|
|
385
|
+
|
|
386
|
+
Galaxy A35 (Exynos 1380 / Mali-G68 MP5)
|
|
387
|
+
CPU (NEON 4-threads): |==================================================| 482.1s
|
|
388
|
+
GPU (Vulkan0): |======| 63.8s [7.56x Faster]
|
|
323
389
|
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
390
|
+
Galaxy S22 (Snapdragon 8 Gen 1 / Adreno 730)
|
|
391
|
+
CPU (NEON 4-threads): |========================================| 394.3s
|
|
392
|
+
GPU (Vulkan0): |=======| 70.9s [5.56x Faster]
|
|
327
393
|
|
|
328
|
-
|
|
394
|
+
Galaxy S20 (Snapdragon 865 / Adreno 650)
|
|
395
|
+
CPU (NEON 4-threads): |==================================================| 512.6s
|
|
396
|
+
GPU (Vulkan0): |============| 128.5s [3.99x Faster]
|
|
329
397
|
```
|
|
330
398
|
|
|
399
|
+
Vulkan native GPU execution achieves a **$4.0\times\text{ to }7.6\times$ throughput speedup** over optimized multi-threaded ARM NEON CPU computation while reducing overall battery drain via the *Race-to-Sleep* principle.
|
|
400
|
+
|
|
331
401
|
---
|
|
332
402
|
|
|
333
403
|
## 8. GPU Interconnect Architecture & Compatibility
|
|
@@ -337,65 +407,65 @@ Termux-STT interfaces directly with Android's Bionic Vulkan loader (`/system/lib
|
|
|
337
407
|
|
|
338
408
|
### 8.2 Silicon Compatibility Matrix
|
|
339
409
|
- **Qualcomm Snapdragon (Adreno 6xx, 7xx, 8xx)**:
|
|
340
|
-
- **Tier-1 Full Support**. Native FP16 compute instructions and high dispatch concurrency deliver RTF performance as fast as **0.
|
|
410
|
+
- **Tier-1 Full Support**. Native FP16 compute instructions and high dispatch concurrency deliver RTF performance as fast as **0.064x** on Snapdragon 8 Elite. Adreno SoftMax workgroups are strictly constrained to hardware subgroup size (64) for rock-solid stability.
|
|
341
411
|
- **Samsung Exynos / MediaTek Dimensity (ARM Mali / Immortalis)**:
|
|
342
|
-
- **Supported**. Mali tile-based architectures benefit from
|
|
412
|
+
- **Supported**. Mali tile-based architectures benefit from continuous compute queues. Full execution across Tiny, Base, Small, and Turbo without shader stalls.
|
|
343
413
|
- **Strict Zero-Silent-Fallback**:
|
|
344
|
-
- Requesting `--device vulkan` without valid Vulkan drivers immediately triggers `PlatformNotSupportedError`, preventing silent fallback to unoptimized CPU execution.
|
|
414
|
+
- Requesting `--device vulkan` without valid Vulkan drivers immediately triggers typed exceptions (`PlatformNotSupportedError`, error code `AMEVA-STT-E002`), preventing silent fallback to unoptimized CPU execution.
|
|
345
415
|
|
|
346
416
|
---
|
|
347
417
|
|
|
348
|
-
##
|
|
418
|
+
## 9. Architectural Trade-Off Analysis & Decision Outcomes
|
|
419
|
+
|
|
420
|
+
In resource-constrained mobile systems engineering, every architectural design represents an explicit compromise between competing physical constraints:
|
|
349
421
|
|
|
350
|
-
|
|
|
351
|
-
| :--- | :--- | :--- | :--- |
|
|
352
|
-
| **
|
|
353
|
-
| **
|
|
354
|
-
| **
|
|
355
|
-
| **
|
|
356
|
-
| **
|
|
422
|
+
| Trade-Off Domain | Physical Constraints & Situation | Architectural Options Evaluated | Decision Executed | Empirical Outcome & Ground Truth |
|
|
423
|
+
| :--- | :--- | :--- | :--- | :--- |
|
|
424
|
+
| **1. Decoding Search Policy** | Autoregressive decoding saturates UMA memory bandwidth and inflates KV cache memory footprint. | **Option A**: Standard Beam Search ($B=5$) prioritizing hypothesis exploration.<br>**Option B**: Single-path Greedy Decoding ($B=1$). | **Selected Option B (Greedy $B=1$)** as production default; preserved Option A as explicit user override flag. | **5x reduction** in decoder memory bus traffic; KV cache shrunk from $249 ext{ MB}$ to $49.8 ext{ MB}$; zero measurable WER degradation on standard benchmark audio; eliminated mobile thermal spikes. |
|
|
425
|
+
| **2. Command Buffer Granularity** | 32-layer Turbo encoder takes $135.5 ext{s}$ per 30s window on Adreno 650, breaching the 270s KGSL TDR on 60s audio. | **Option A**: Fragment graph into 32-node sub-batches (`GGML_VK_MAX_NODES_PER_SUBMIT=32`).<br>**Option B**: Retain monolithic submission and enforce Fail-Fast. | **Selected Option B (Monolithic)** for production default; rejected forced fragmentation. | Subdividing submissions overloaded Qualcomm's 2021 kernel ringbuffer, triggering `IOCTL_KGSL_GPU_COMMAND: errno 35 (EDEADLK)` at $136.7 ext{s}$. Monolithic execution allows S22/S25/A35/A53 to achieve peak throughput without synchronization pipeline stalls. |
|
|
426
|
+
| **3. Heterogeneous Execution Routing** | High encoder latency suggested splitting encoder to GPU and decoder to CPU. | **Option A**: Dynamic asymmetric hybrid offloading across CPU/GPU.<br>**Option B**: Pure native Vulkan ABI pipeline binding. | **Selected Option B (Pure Native ABI)**; permanently blacklisted asymmetric offloading. | Avoided $7.68 ext{ MB}$ per layer inter-device tensor ping-pong over non-coherent UMA caches; eliminated thread synchronization jitter and delivered a unified, maintainable codebase. |
|
|
427
|
+
| **4. Hardware Support Boundaries** | Snapdragon 865 cannot complete 60s monolithic Turbo without TDR due to 2021 driver limitations. | **Option A**: Hardcode software check blocking Turbo model selection on S20.<br>**Option B**: Zero artificial restrictions, allow users complete freedom, let kernel return native exit codes. | **Selected Option B (Zero Artificial Blocks)** with full documentation transparency. | Preserves OpenSSF open-source compliance; allows users running shorter audio clips ($<30 ext{s}$) or custom kernels to execute Turbo unimpeded; maintains complete engineering honesty. |
|
|
428
|
+
| **5. Virtual Memory on 6GB SoCs** | Exynos 1280 (A53) has 6GB physical RAM; loading 809M Turbo caused severe swap thrashing ($>900 ext{s}$ timeout). | **Option A**: Restrict A53 to Small models only.<br>**Option B**: Provision +2GB zRAM swap backing store (RAM Plus) accepting minor compression overhead. | **Selected Option B (+2GB RAM Plus)** in operating system settings. | Reduced page fault rate $P_{fault} < 0.0001$; eliminated `kswapd0` CPU spin loops; Small finished in **$155.83 ext{s}$** and Turbo finished in **$759.52 ext{s}$** with $100\%$ transcript fidelity. |
|
|
357
429
|
|
|
358
|
-
|
|
430
|
+
> For the comprehensive academic analysis, kernel watchdog expiry mathematical proofs, and driver boundary traces, refer to the [Master Technical Treatise (English)](docs/research/on_device_vulkan_stt_master_treatise.md) and [연구 백서 (Korean)](docs/research/on_device_vulkan_stt_master_treatise_kor.md).
|
|
359
431
|
|
|
360
432
|
---
|
|
361
433
|
|
|
362
|
-
##
|
|
434
|
+
## 10. Hardware Requirements & Operational Limits
|
|
363
435
|
|
|
364
|
-
###
|
|
436
|
+
### 10.1 Hardware Specifications
|
|
365
437
|
|
|
366
438
|
| Specification Metric | Minimum Requirements | Recommended Production Spec |
|
|
367
439
|
| :--- | :--- | :--- |
|
|
368
440
|
| **Operating System** | Android 9.0+ (API level 28+) / Linux 5.4+ | Android 12.0+ (API level 31+) |
|
|
369
441
|
| **Architecture** | ARM64 (aarch64) or x86_64 | ARM64-v8a / v9a |
|
|
370
|
-
| **System RAM** | 2 GB Total Unified RAM | 4 GB+ Unified RAM |
|
|
442
|
+
| **System RAM** | 2 GB Total Unified RAM | 4 GB+ Unified RAM (6GB+ for Turbo) |
|
|
371
443
|
| **Storage Footprint** | 150 MB (Vosk) / 300 MB (Whisper Base) | 1 GB Free Flash Storage |
|
|
372
444
|
| **Audio Subsystem** | Termux-API Microphone Permissions | 16kHz PCM Audio Capture Support |
|
|
373
445
|
|
|
374
|
-
###
|
|
446
|
+
### 10.2 Operational Limits & Best Practices
|
|
375
447
|
- **32-Bit ARM (armeabi-v7a)**: Not supported for Whisper GPU neural inference. Use Vosk CFFI for legacy 32-bit hardware.
|
|
376
448
|
- **Microphone Permissions**: Real-time microphone listening (`termux-stt listen`) requires Android microphone permission granted to Termux: `termux-microphone-record`.
|
|
449
|
+
- **6GB RAM Mid-Range Devices**: For large models (Turbo 809M) on 6GB devices such as Galaxy A53, enable +2GB RAM Plus (zRAM) in Android settings to eliminate page thrashing.
|
|
377
450
|
|
|
378
451
|
---
|
|
379
452
|
|
|
380
|
-
##
|
|
453
|
+
## 11. 24/7 Unattended Background Execution Guide
|
|
381
454
|
|
|
382
|
-
Android aggressively kills background user-space processes inside Termux unless battery and process monitor policies are explicitly configured
|
|
455
|
+
Android aggressively kills background user-space processes inside Termux unless battery and process monitor policies are explicitly configured:
|
|
383
456
|
|
|
384
|
-
###
|
|
457
|
+
### 11.1 Stage 1: Termux Kernel Wake-Lock
|
|
385
458
|
Prevent the mobile CPU from entering low-power sleep states:
|
|
386
459
|
```bash
|
|
387
|
-
# Acquire persistent CPU wake-lock
|
|
388
460
|
termux-wake-lock
|
|
389
461
|
```
|
|
390
462
|
|
|
391
|
-
###
|
|
463
|
+
### 11.2 Stage 2: Android GUI Battery Optimization Exemption
|
|
392
464
|
1. Open **Android Settings > Apps > Termux > Battery**.
|
|
393
465
|
2. Set battery policy to **Unrestricted** (Disable power-saving restrictions).
|
|
394
466
|
3. Under **Permissions**, grant **Microphone** and **Display over other apps**.
|
|
395
467
|
|
|
396
|
-
###
|
|
397
|
-
Android 12+ terminates background processes exceeding child process thresholds. Execute these commands via ADB:
|
|
398
|
-
|
|
468
|
+
### 11.3 Stage 3: ADB Phantom Process Killer Exemption (Android 12+)
|
|
399
469
|
```bash
|
|
400
470
|
# Disable Android Phantom Process Killer
|
|
401
471
|
adb shell device_config put activity_manager max_phantom_processes 2147483647
|
|
@@ -408,7 +478,7 @@ adb shell settings get global settings_enable_monitor_phantom_procs
|
|
|
408
478
|
|
|
409
479
|
---
|
|
410
480
|
|
|
411
|
-
##
|
|
481
|
+
## 12. Open Source License
|
|
412
482
|
|
|
413
483
|
Termux-STT is open-sourced under the **MIT License**.
|
|
414
484
|
|
|
@@ -436,7 +506,7 @@ SOFTWARE.
|
|
|
436
506
|
|
|
437
507
|
---
|
|
438
508
|
|
|
439
|
-
##
|
|
509
|
+
## 13. SEO Technical Keywords & Ecosystem Metadata
|
|
440
510
|
|
|
441
511
|
`termux`, `stt`, `speech-to-text`, `whisper`, `whisper-cpp`, `vosk`, `sherpa-onnx`, `diarization`, `speaker-diarization`, `voice-recognition`, `audio-transcription`, `on-device-ai`, `edge-ai`, `mobile-ai`, `vulkan`, `vulkan-compute`, `gpu-acceleration`, `real-time-factor`, `low-latency`, `arm64`, `android`, `snapdragon`, `adreno`, `exynos`, `arm-mali`, `zero-compilation`, `vad`, `voice-activity-detection`, `silero-vad`, `x-vector`, `k-means`, `clustering`, `srt-export`, `vtt-export`, `rttm`, `microphone-streaming`, `headless-audio`, `pulseaudio`, `offline-speech`, `privacy-first`, `termux-aichain`, `termux-tts`, `termux-llamacpp`, `termux-diffusion`, `ameva-runtime`, `ggml`, `quantization`, `bionic-libc`, `autonomous-agents`, `voice-assistant`
|
|
442
512
|
|
|
@@ -446,3 +516,5 @@ SOFTWARE.
|
|
|
446
516
|
- **Official Documentation Portal**: [https://uno-km.github.io/termux-stt/](https://uno-km.github.io/termux-stt/)
|
|
447
517
|
- **GitHub Repository**: [https://github.com/uno-km/termux-stt](https://github.com/uno-km/termux-stt)
|
|
448
518
|
- **AMEVA Foundation Portal**: [https://uno-km.vercel.app/foundation/index.html](https://uno-km.vercel.app/foundation/index.html)
|
|
519
|
+
- **Master Research Treatise (English)**: [docs/research/on_device_vulkan_stt_master_treatise.md](docs/research/on_device_vulkan_stt_master_treatise.md)
|
|
520
|
+
- **연구 백서 (Korean)**: [docs/research/on_device_vulkan_stt_master_treatise_kor.md](docs/research/on_device_vulkan_stt_master_treatise_kor.md)
|
package/package.json
CHANGED