termux-stt 1.1.5 → 1.1.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +28 -350
- package/index.js +0 -14
- package/lib/whisper.js +17 -1
- package/package.json +1 -4
- package/README.pypi.md +0 -34
package/README.md
CHANGED
|
@@ -1,389 +1,67 @@
|
|
|
1
|
-
#
|
|
1
|
+
# termux-stt
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
```
|
|
6
|
-
████████╗███████╗██████╗ ███╗ ███╗██╗ ██╗██╗ ██╗ ███████╗████████╗████████╗
|
|
7
|
-
╚══██╔══╝██╔════╝██╔══██╗████╗ ████║██║ ██║╚██╗██╔╝ ██╔════╝╚══██╔══╝╚══██╔══╝
|
|
8
|
-
██║ █████╗ ██████╔╝██╔████╔██║██║ ██║ ╚███╔╝ █████╗███████╗ ██║ ██║
|
|
9
|
-
██║ ██╔══╝ ██╔══██╗██║╚██╔╝██║██║ ██║ ██╔██╗ ╚════╝╚════██║ ██║ ██║
|
|
10
|
-
██║ ███████╗██║ ██║██║ ╚═╝ ██║╚██████╔╝██╔╝ ██╗ ███████║ ██║ ██║
|
|
11
|
-
╚═╝ ╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝ ╚═╝ ╚═╝
|
|
12
|
-
```
|
|
13
|
-
|
|
14
|
-
**Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux**
|
|
15
|
-
*Dual-Engine Architecture (Python & Node.js / TypeScript) with Native Bionic ARM64 Acceleration & 0 PyTorch Dependency*
|
|
3
|
+
> **Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Node.js on Android Termux.**
|
|
16
4
|
|
|
17
5
|
<p align="center">
|
|
18
|
-
<a href="https://
|
|
19
|
-
<a href="https://
|
|
20
|
-
<a href="https://
|
|
21
|
-
<
|
|
6
|
+
<a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/npm/v/termux-stt.svg?style=flat-square&color=cb3837&logo=npm&logoColor=white" alt="npm Version" /></a>
|
|
7
|
+
<a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/badge/npm%20Downloads-active-cb3837?style=flat-square&logo=npm&logoColor=white" alt="npm Downloads" /></a>
|
|
8
|
+
<a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg?style=flat-square" alt="License" /></a>
|
|
9
|
+
<img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
|
|
22
10
|
</p>
|
|
23
11
|
|
|
24
|
-
<p align="center">
|
|
25
|
-
<a href="https://uno-km.vercel.app/lib/stt/"><img src="https://img.shields.io/badge/Official_Docs-uno--km.vercel.app%2Flib%2Fstt-004499?style=for-the-badge&logo=vercel&logoColor=white" alt="Live Docs" /></a>
|
|
26
|
-
<a href="https://uno-km.vercel.app/lib/stt/demo.html"><img src="https://img.shields.io/badge/Live_Showcase-▶_Audio_Player-00f5d4?style=for-the-badge&logo=googlechrome&logoColor=0b132b" alt="Live Audio Showcase" /></a>
|
|
27
|
-
<a href="https://github.com/uno-km/termux-stt"><img src="https://img.shields.io/github/stars/uno-km/termux-stt?style=for-the-badge&color=gold&logo=github" alt="GitHub Stars" /></a>
|
|
28
|
-
<a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg?style=for-the-badge" alt="License" /></a>
|
|
29
|
-
</p>
|
|
30
|
-
|
|
31
|
-
<p align="center">
|
|
32
|
-
<img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64%2Faarch64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
|
|
33
|
-
<img src="https://img.shields.io/badge/Engines-whisper.cpp%20%7C%20Vosk%20%7C%20Sherpa--ONNX-38bdf8?style=flat-square" alt="Engines" />
|
|
34
|
-
<img src="https://img.shields.io/badge/Diarization-128d%20X--Vector%20+%20Pure%20Python-a855f7?style=flat-square" alt="Diarization" />
|
|
35
|
-
<img src="https://img.shields.io/badge/RAM-Under%20350MB%20(Tiny%2FBase)-10b981?style=flat-square&logo=shield&logoColor=white" alt="RAM" />
|
|
36
|
-
<img src="https://img.shields.io/badge/Foundation-AOSF_Tier_1-orange?style=flat-square" alt="Foundation" />
|
|
37
|
-
</p>
|
|
38
|
-
|
|
39
|
-
<br/>
|
|
40
|
-
|
|
41
|
-
**[Live Audio Showcase & Demo](https://uno-km.vercel.app/lib/stt/demo.html)** • **[Official Documentation Site (13 Languages)](https://uno-km.vercel.app/lib/stt/)** • **[AMEVA Foundation](https://uno-km.vercel.app/docs/foundation/)** • **[Quickstart](#1-quick-scenario-playbook)** • **[Architecture](#2-why-termux-stt-architectural-pillars)** • **[Benchmarks](#3-empirical-benchmarks-galaxy-a35--exynos-1380)**
|
|
42
|
-
|
|
43
|
-
</div>
|
|
44
|
-
|
|
45
12
|
---
|
|
46
13
|
|
|
47
|
-
##
|
|
48
|
-
|
|
49
|
-
> **"$0 Cloud Egress, 0% External Data Leaks. Transforming every smartphone into an independent on-device AI workstation."**
|
|
50
|
-
> The **AMEVA Open-Source Foundation (AOSF)** builds next-generation, client-centric AI runtimes spanning on-device large models, browser automation, neural network training, and speaker diarization.
|
|
51
|
-
|
|
52
|
-
<div align="center">
|
|
14
|
+
## 1. Quick Start
|
|
53
15
|
|
|
54
|
-
|
|
55
|
-
| :--- | :--- | :--- | :---: |
|
|
56
|
-
| 🎙️ **[termux-stt](https://github.com/uno-km/termux-stt)** | [](https://opencollective.com/ameva-fund) [](https://github.com/sponsors/uno-km)<br/>[](https://pypi.org/project/termux-stt/) [](https://www.npmjs.com/package/termux-stt) | **Integrated On-Device STT & Pure Python 128d X-Vector Diarization** (Whisper + Vosk + Sherpa) | **[Showcase](https://uno-km.github.io/termux-stt/showcase.html)** • **[Docs](https://uno-km.vercel.app/lib/stt/)** |
|
|
57
|
-
| 🎨 **[termux-diffusion](https://github.com/uno-km/termux-diffusion)** | [](https://pypi.org/project/termux-diffusion/) [](https://www.npmjs.com/package/termux-diffusion) | **Mobile On-Device Stable Diffusion Image Generation** (bfloat16 ARM NEON acceleration) | **[Docs](https://uno-km.github.io/termux-diffusion/)** |
|
|
58
|
-
| 🌐 **[termux-playwright](https://github.com/uno-km/termux-playwright)** | [](https://pypi.org/project/termux-playwright/) [](https://www.npmjs.com/package/termux-playwright) | **Non-Root Native Headless Chromium Browser Automation & Scraping** | **[Docs](https://uno-km.github.io/termux-playwright/)** |
|
|
59
|
-
| 🧠 **[termux-train](https://github.com/uno-km/termux-train)** | [](https://pypi.org/project/termux-train/) | **Mobile Native Autograd Neural Network Training & LoRA Fine-Tuning** | **[Docs](https://uno-km.vercel.app/lib/train/)** |
|
|
60
|
-
| 🖥️ **[AMEVA Workstation](https://github.com/uno-km/AMEVA-Workstation-Web)** | [](https://ameva-workstation-web-core.vercel.app/) | **100% Client-Side WebGPU Multimodal Document Intelligence Workspace** | **[Live Demo](https://ameva-workstation-web-core.vercel.app/)** |
|
|
61
|
-
| ⚡ **[AMEVA-Forge](https://github.com/uno-km/ameva-forge)** | [](https://uno-km.github.io/ameva-forge/demo.html) | **Real-Time 3D Neural Studio & WebGPU Visualization Engine** | **[Live Demo](https://uno-km.github.io/ameva-forge/demo.html)** |
|
|
16
|
+
### 1.1 Installation
|
|
62
17
|
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
---
|
|
66
|
-
|
|
67
|
-
## 1. Installation & Provisioning
|
|
68
|
-
|
|
69
|
-
### Method A: 1-Click Automated Installation (Recommended)
|
|
70
|
-
`termux-stt install` automatically provisions lightweight system dependencies (`ffmpeg`, `libbluray`, `libxml2`) and fetches the **pre-compiled Bionic ARM64 `whisper-cli` binary from GitHub Releases in <1s** (with automatic local Clang/NEON compilation fallback if offline).
|
|
71
|
-
|
|
72
|
-
#### Python SDK:
|
|
73
|
-
```bash
|
|
74
|
-
pkg update -y && pkg install python ffmpeg git -y
|
|
75
|
-
pip install --upgrade termux-stt && termux-stt install
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
#### Node.js / TypeScript:
|
|
18
|
+
#### 1-Click Automated Setup:
|
|
79
19
|
```bash
|
|
80
20
|
pkg update -y && pkg install nodejs-lts ffmpeg git -y
|
|
81
|
-
npm install -g termux-stt
|
|
21
|
+
npm install -g termux-stt
|
|
22
|
+
termux-stt install
|
|
82
23
|
```
|
|
83
24
|
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
### Method B: Manual Pre-Compiled Binary Download (Direct GitHub Release)
|
|
87
|
-
If you prefer to manually download the pre-compiled ARM64 native binary without building from source:
|
|
88
|
-
|
|
25
|
+
#### Manual Pre-Compiled Binary Download (Direct GitHub Release):
|
|
89
26
|
```bash
|
|
90
|
-
# 1. Download official ARM64 Bionic binary directly into ~/.local/bin
|
|
91
27
|
mkdir -p ~/.local/bin
|
|
92
28
|
curl -sL "https://github.com/uno-km/termux-stt/releases/download/v1.1.1/whisper-cli-arm64-android" -o ~/.local/bin/whisper-cli
|
|
93
29
|
chmod 755 ~/.local/bin/whisper-cli
|
|
94
30
|
ln -sf ~/.local/bin/whisper-cli ~/.local/bin/whisper-cpp
|
|
95
|
-
|
|
96
|
-
# 2. Ensure PATH includes ~/.local/bin
|
|
97
31
|
export PATH=$HOME/.local/bin:$PATH
|
|
98
|
-
|
|
99
|
-
# 3. Verify installation
|
|
100
|
-
whisper-cli --help
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
---
|
|
104
|
-
|
|
105
|
-
### Method C: Manual Source Compilation (From Scratch)
|
|
106
|
-
If you wish to compile `whisper.cpp` directly with Snapdragon / Cortex NEON vector optimizations:
|
|
107
|
-
|
|
108
|
-
```bash
|
|
109
|
-
pkg install -y clang cmake make git ffmpeg
|
|
110
|
-
git clone --depth 1 https://github.com/ggerganov/whisper.cpp.git ~/tmp/whisper.cpp
|
|
111
|
-
cmake -B ~/tmp/whisper.cpp/build -S ~/tmp/whisper.cpp -DWHISPER_NEON=ON -DCMAKE_BUILD_TYPE=Release
|
|
112
|
-
cmake --build ~/tmp/whisper.cpp/build -j$(nproc)
|
|
113
|
-
cp ~/tmp/whisper.cpp/build/bin/whisper-cli ~/.local/bin/whisper-cli
|
|
114
|
-
chmod 755 ~/.local/bin/whisper-cli
|
|
115
32
|
```
|
|
116
33
|
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
### 🎙️ Try with the Included Sample Audio! (JFK 1-Minute Speech)
|
|
34
|
+
### 1.2 CLI Usage & Instant Sample Test
|
|
120
35
|
|
|
121
|
-
`termux-stt`
|
|
36
|
+
`termux-stt` includes a 60-second JFK inaugural address speech (`samples/jfk_1min.wav`) for immediate testing:
|
|
122
37
|
|
|
123
|
-
#### Option A: CLI One-Liner (Python / npm)
|
|
124
38
|
```bash
|
|
125
|
-
# 1.
|
|
126
|
-
|
|
39
|
+
# 1. Transcribe sample speech with Whisper Tiny (Generates SRT subtitles)
|
|
40
|
+
termux-stt transcribe samples/jfk_1min.wav --engine whisper --model tiny --format srt
|
|
127
41
|
|
|
128
|
-
# 2. Transcribe
|
|
129
|
-
termux-stt
|
|
130
|
-
|
|
131
|
-
# 3. Transcribe and separate speakers (128d X-Vector Diarization)
|
|
132
|
-
termux-stt diarize jfk_1min.wav --speakers 2 --format text
|
|
133
|
-
|
|
134
|
-
# 4. Benchmark on-device latency & RTF
|
|
135
|
-
termux-stt benchmark --audio jfk_1min.wav --model tiny
|
|
42
|
+
# 2. Transcribe with 128d Pure Python Speaker Diarization
|
|
43
|
+
termux-stt diarize samples/jfk_1min.wav --speakers 2
|
|
136
44
|
```
|
|
137
45
|
|
|
138
|
-
|
|
139
|
-
```python
|
|
140
|
-
from termux_stt import create_engine
|
|
141
|
-
|
|
142
|
-
# 1. Initialize Whisper Engine (auto-loads native ARM NEON binary)
|
|
143
|
-
engine = create_engine("whisper", model="tiny", lang="en", threads=4)
|
|
144
|
-
|
|
145
|
-
# 2. Transcribe JFK 60s speech
|
|
146
|
-
result = engine.transcribe("samples/jfk_1min.wav")
|
|
46
|
+
---
|
|
147
47
|
|
|
148
|
-
|
|
149
|
-
print("SRT Subtitles:\n", result.to_srt())
|
|
150
|
-
```
|
|
48
|
+
## 2. Programmatic Node.js API
|
|
151
49
|
|
|
152
|
-
#### Option C: Node.js / TypeScript
|
|
153
50
|
```javascript
|
|
154
|
-
const { createEngine } = require(
|
|
51
|
+
const { createEngine } = require('termux-stt');
|
|
155
52
|
|
|
156
53
|
async function main() {
|
|
157
|
-
const engine = createEngine(
|
|
158
|
-
const result = await engine.transcribe(
|
|
159
|
-
|
|
160
|
-
console.log(
|
|
161
|
-
console.log("SRT Subtitles:\n", result.toSrt());
|
|
54
|
+
const engine = createEngine('whisper', { model: 'tiny', lang: 'en', threads: 4 });
|
|
55
|
+
const result = await engine.transcribe('samples/jfk_1min.wav');
|
|
56
|
+
console.log('Transcription:', result.text);
|
|
57
|
+
console.log('SRT Subtitles:\n', result.toSrt());
|
|
162
58
|
}
|
|
163
|
-
main();
|
|
164
|
-
```
|
|
165
|
-
|
|
166
|
-
---
|
|
167
|
-
|
|
168
|
-
## 2. Advanced Scenarios
|
|
169
|
-
|
|
170
|
-
### [Streaming] Scenario 3: Real-Time Microphone Streaming
|
|
171
|
-
|
|
172
|
-
Speak into your smartphone microphone and receive real-time transcribed text with sub-second latency:
|
|
173
|
-
|
|
174
|
-
```python
|
|
175
|
-
from termux_stt import create_engine
|
|
176
|
-
|
|
177
|
-
# Use ultra-lightweight Tiny model (RTF 0.80 on Exynos 1380)
|
|
178
|
-
engine = create_engine("whisper", model="tiny", lang="ko")
|
|
179
|
-
|
|
180
|
-
print("🎙️ Listening... Speak into your phone microphone (Ctrl+C to stop)")
|
|
181
|
-
for segment in engine.stream_mic():
|
|
182
|
-
print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")
|
|
183
|
-
```
|
|
184
|
-
|
|
185
|
-
---
|
|
186
|
-
|
|
187
|
-
### [Diarize] Scenario 4: Hybrid Speaker Diarization ("Who Spoke When?")
|
|
188
|
-
|
|
189
|
-
Run high-precision speaker diarization without PyTorch or CUDA:
|
|
190
59
|
|
|
191
|
-
|
|
192
|
-
from termux_stt import create_engine
|
|
193
|
-
|
|
194
|
-
# Hybrid Pipeline: Vosk 128d X-Vector + Whisper STT + Pure Python K-Means
|
|
195
|
-
engine = create_engine("hybrid", lang="ko", num_speakers=2)
|
|
196
|
-
result = engine.diarize("interview.wav")
|
|
197
|
-
|
|
198
|
-
for seg in result.segments:
|
|
199
|
-
print(f"[{seg.speaker}] ({seg.start:.1f}s - {seg.end:.1f}s): {seg.text}")
|
|
200
|
-
```
|
|
201
|
-
|
|
202
|
-
*Output Example:*
|
|
203
|
-
```text
|
|
204
|
-
[Speaker_0] (0.0s - 3.5s): Welcome to today's financial intelligence briefing.
|
|
205
|
-
[Speaker_1] (3.8s - 7.2s): Major equity indices closed higher following sustained foreign institutional inflows.
|
|
206
|
-
[Speaker_0] (7.5s - 10.1s): What is the latest outlook on the semiconductor sector?
|
|
207
|
-
```
|
|
208
|
-
|
|
209
|
-
---
|
|
210
|
-
|
|
211
|
-
## 2. 🏛️ Why termux-stt? Architectural Pillars
|
|
212
|
-
|
|
213
|
-
```
|
|
214
|
-
┌─────────────────────────────────────────────────────────────────────────────┐
|
|
215
|
-
│ termux-stt Architecture │
|
|
216
|
-
├────────────────────────┬────────────────────────┬───────────────────────────┤
|
|
217
|
-
│ User Interface Layer │ Engine Abstraction │ Output / Export Layer │
|
|
218
|
-
│ • Python API │ • EngineRegistry │ • JSON / SRT / VTT │
|
|
219
|
-
│ • Node.js API │ • create_engine() │ • RTTM (Diarization) │
|
|
220
|
-
│ • CLI (termux-stt) │ • ModelHub & Cache │ • Streaming Callback │
|
|
221
|
-
├────────────────────────┼────────────────────────┼───────────────────────────┤
|
|
222
|
-
│ Core Pipeline Layer │
|
|
223
|
-
│ Audio Loader (7 Formats) ➔ Preprocessor (16kHz Mono) ➔ Silero-VAD Filter │
|
|
224
|
-
│ ➔ Multi-Engine STT (whisper.cpp / Vosk / Sherpa-ONNX) │
|
|
225
|
-
│ ➔ Hybrid Diarizer (128d X-Vector ➔ Pure Python Cosine/K-Means ➔ Time Align)│
|
|
226
|
-
├─────────────────────────────────────────────────────────────────────────────┤
|
|
227
|
-
│ Platform & Process Isolation │
|
|
228
|
-
│ • Subprocess Isolation (Host Python never crashes on C++ Segfault) │
|
|
229
|
-
│ • MobileGuard (WakeLock, Doze Mode Bypass, Phantom Process Killer Shield) │
|
|
230
|
-
│ • Bionic ARM64 NEON & FP16 SIMD Acceleration │
|
|
231
|
-
└─────────────────────────────────────────────────────────────────────────────┘
|
|
232
|
-
```
|
|
233
|
-
|
|
234
|
-
1. **Subprocess Isolation**: C++ inference runs in isolated native subprocesses. Memory errors or Segfaults never kill the host Python/Node.js application.
|
|
235
|
-
2. **Pure Python Clustering**: Cosine distance matrix and K-Means clustering are implemented with 0 external dependencies (`numpy` / `scikit-learn` optional, not required).
|
|
236
|
-
3. **Automated Android Bionic Fixes**: Solves `sys.platform = 'linux'` spoofing, libvosk CFFI extraction, and ffmpeg audio format normalization (`16kHz / 1ch / PCM s16le`) under the hood.
|
|
237
|
-
4. **Mobile Battery & CPU Guard**: Manages Android WakeLocks and background task states to prevent process termination when the screen locks.
|
|
238
|
-
|
|
239
|
-
---
|
|
240
|
-
|
|
241
|
-
## 3. 📊 Empirical Benchmarks (Galaxy A35 / Exynos 1380)
|
|
242
|
-
|
|
243
|
-
> *Measured on Samsung Galaxy A35 5G (Exynos 1380 4x A78 + 4x A55, 6GB RAM, Android 14 Termux).*
|
|
244
|
-
|
|
245
|
-
| Engine / Pipeline | Model | Peak RAM | RTF (Speed) | KO Accuracy | Diarization Support | Termux Rating |
|
|
246
|
-
| :--- | :--- | :---: | :---: | :---: | :---: | :---: |
|
|
247
|
-
| **whisper.cpp** | `ggml-tiny` (39M) | **~150 MB** | **0.80** | 85% | ❌ External | ⭐⭐⭐⭐⭐ |
|
|
248
|
-
| **whisper.cpp** | `ggml-base` (74M) | ~250 MB | 1.20 | 88% | ❌ External | ⭐⭐⭐⭐ |
|
|
249
|
-
| **whisper.cpp** | `ggml-medium` (769M) | ~1.5 GB | 3.40 | **95%+** | ❌ External | ⭐⭐⭐ (Golden Acc) |
|
|
250
|
-
| **Vosk** | `small-ko-0.22` (42M) | **~100 MB** | **0.25** | 78% | ✅ 128d X-Vector | ⭐⭐⭐ |
|
|
251
|
-
| **Sherpa-ONNX** | `Zipformer` | ~300 MB | **0.42** | 86% | ✅ CAM++ | ⭐⭐⭐⭐ |
|
|
252
|
-
| **Pyannote.audio 3.1** | `diarization-3.1` | **> 3.5 GB** | 2.80~3.50 | N/A | ✅ Gold Standard | ❌ OOM Crashes |
|
|
253
|
-
| **termux-stt (Hybrid)** | `Vosk + Whisper Base`| **~350 MB** | **1.45** | **92%+** | **✅ Built-in K-Means** | **⭐⭐⭐⭐⭐ (Recommended)** |
|
|
254
|
-
|
|
255
|
-
---
|
|
256
|
-
|
|
257
|
-
## 4. ⚙️ Engine Comparison Matrix
|
|
258
|
-
|
|
259
|
-
| Feature | `whisper.cpp` | `Vosk` | `Sherpa-ONNX` | `Hybrid (Vosk+Whisper)` |
|
|
260
|
-
| :--- | :---: | :---: | :---: | :---: |
|
|
261
|
-
| **Primary Strength** | Highest Text Accuracy | Ultra-Low RAM & Fast | Ultra-Low Latency | **STT + Diarization Combined** |
|
|
262
|
-
| **Memory Footprint** | 150MB ~ 1.5GB | **< 100MB** | 300MB ~ 500MB | **~ 350MB** |
|
|
263
|
-
| **Real-Time Factor (RTF)** | 0.80 (Tiny) | **0.25 (Blazing)** | **0.42 (Fast)** | 1.45 (Full Pipeline) |
|
|
264
|
-
| **Speaker Diarization** | ❌ None | ⚠️ Basic X-Vector | ⚠️ CAM++ C++ | **✅ High-Precision Aligned** |
|
|
265
|
-
| **Recommended Use Case** | Quality Transcripts | Embedded / Low Spec | Live Voice Assistant | **Meetings / Interviews** |
|
|
266
|
-
|
|
267
|
-
---
|
|
268
|
-
|
|
269
|
-
## 5. 📚 Complete API Reference Summary
|
|
270
|
-
|
|
271
|
-
### Python API
|
|
272
|
-
|
|
273
|
-
```python
|
|
274
|
-
import termux_stt
|
|
275
|
-
|
|
276
|
-
# Create Engine with fine-grained control parameters
|
|
277
|
-
engine = termux_stt.create_engine(
|
|
278
|
-
engine="whisper", # "whisper" | "vosk" | "sherpa" | "hybrid"
|
|
279
|
-
model="base", # "tiny" | "base" | "small" | "medium" | "custom"
|
|
280
|
-
lang="ko", # ISO 639-1 language code
|
|
281
|
-
num_speakers=0, # 0 = disabled, 2+ = enable diarization
|
|
282
|
-
threads=4, # CPU threads (defaults to big cores count)
|
|
283
|
-
vad=True, # Enable Silero-VAD silence stripping
|
|
284
|
-
quantization="q5_1", # "f16" | "q8_0" | "q5_1" | "q4_0"
|
|
285
|
-
prompt="Financial briefing", # Initial decoding context / vocabulary
|
|
286
|
-
beam_size=5, # Beam search beam size
|
|
287
|
-
temperature=0.0 # Sampling temperature
|
|
288
|
-
)
|
|
289
|
-
|
|
290
|
-
# Transcribe File
|
|
291
|
-
result = engine.transcribe("audio.wav")
|
|
292
|
-
# Returns: TranscriptResult(text=str, segments=List[Segment], language=str, duration=float)
|
|
293
|
-
|
|
294
|
-
# Export Methods
|
|
295
|
-
result.to_json() # Structured JSON string
|
|
296
|
-
result.to_srt() # Standard SRT subtitle format
|
|
297
|
-
result.to_vtt() # WebVTT subtitle format
|
|
298
|
-
result.to_rttm() # NIST RTTM diarization format
|
|
299
|
-
|
|
300
|
-
# Stream Microphone
|
|
301
|
-
for seg in engine.stream_mic(duration=30.0):
|
|
302
|
-
print(f"[{seg.speaker}] {seg.text}")
|
|
303
|
-
|
|
304
|
-
# Speaker Diarization
|
|
305
|
-
diar_result = engine.diarize("meeting.wav", num_speakers=2)
|
|
306
|
-
```
|
|
307
|
-
|
|
308
|
-
### CLI Reference
|
|
309
|
-
|
|
310
|
-
```bash
|
|
311
|
-
# General Syntax
|
|
312
|
-
termux-stt [COMMAND] [OPTIONS] [FILE]
|
|
313
|
-
|
|
314
|
-
# Commands
|
|
315
|
-
termux-stt transcribe [FILE] # Transcribe audio file (--prompt, --beam-size, --translate)
|
|
316
|
-
termux-stt listen # Real-time microphone transcription
|
|
317
|
-
termux-stt diarize [FILE] # Perform speaker diarization
|
|
318
|
-
termux-stt models list # List installed and available models
|
|
319
|
-
termux-stt models download [M] # Download specific model
|
|
320
|
-
termux-stt doctor # Run hardware and environment diagnostics
|
|
321
|
-
termux-stt benchmark # Run performance benchmark suite
|
|
60
|
+
main();
|
|
322
61
|
```
|
|
323
62
|
|
|
324
63
|
---
|
|
325
64
|
|
|
326
|
-
##
|
|
327
|
-
|
|
328
|
-
### Q1: `pip install vosk` fails with CMake or wheel error on Android
|
|
329
|
-
* **Cause**: Vosk does not publish official prebuilt aarch64-android wheels on PyPI.
|
|
330
|
-
* **Solution**: `termux-stt-install` automatically extracts `libvosk.so` from the official Android AAR and generates the CFFI bindings.
|
|
331
|
-
|
|
332
|
-
### Q2: Whisper crashes on 44.1kHz stereo MP3/M4A files
|
|
333
|
-
* **Cause**: Whisper models strictly require single-channel 16,000Hz 16-bit PCM WAV.
|
|
334
|
-
* **Solution**: `termux-stt` automatically runs `ffmpeg` normalization on any audio format (`mp3`, `m4a`, `flac`, `ogg`, `opus`, `webm`).
|
|
335
|
-
|
|
336
|
-
### Q3: Process killed after 10 minutes in background
|
|
337
|
-
* **Cause**: Android Phantom Process Killer terminates background tasks.
|
|
338
|
-
* **Solution**: Enable Termux WakeLock (`termux-wake-lock`) and disable battery optimization for Termux in Android Settings.
|
|
339
|
-
|
|
340
|
-
---
|
|
341
|
-
|
|
342
|
-
## 7. 🔍 15-Part Empirical Research Blog Series
|
|
343
|
-
|
|
344
|
-
This framework is built upon the exhaustive 15-part research series published on [Eunho Kim's Technical Engineering Blog](https://uno-kim.tistory.com/):
|
|
345
|
-
|
|
346
|
-
1. [[Whisper.cpp] #1. Edge Agent AI: Whisper.cpp Speech Processing (Base vs Tiny)](https://uno-kim.tistory.com/467)
|
|
347
|
-
2. [[Audio Extraction] #2. Extracting Specific Audio Segments on Android](https://uno-kim.tistory.com/468)
|
|
348
|
-
3. [[Whisper.cpp] #3. Korean STT Conversion Comparison (4 Models)](https://uno-kim.tistory.com/469)
|
|
349
|
-
4. [[Comparison-1] #4. STT + Speaker Diarization: 3 Lightweight Engines + Pyannote](https://uno-kim.tistory.com/472)
|
|
350
|
-
5. [[Comparison-2] #5. Sherpa-ONNX Execution, Diarization, and Troubleshooting](https://uno-kim.tistory.com/473)
|
|
351
|
-
6. [[Comparison-3] #6. Speaker Diarization using Pyannote Model](https://uno-kim.tistory.com/471)
|
|
352
|
-
7. [[Pyannote] Troubleshooting Diarization in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/470)
|
|
353
|
-
8. [[Comparison-4] #7. Vosk Execution, Speaker Diarization, and Troubleshooting](https://uno-kim.tistory.com/475)
|
|
354
|
-
9. [[Vosk] Troubleshooting Vosk in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/474)
|
|
355
|
-
10. [[Comparison-5] #8. Vosk / Pyannote / Sherpa-ONNX / Whisper.cpp Comprehensive Comparison](https://uno-kim.tistory.com/476)
|
|
356
|
-
11. [[Comparison-6] #9. Vosk + Whisper.cpp Hybrid Pipeline & X-Vector Diarization](https://uno-kim.tistory.com/477)
|
|
357
|
-
12. [[Development-1] #10. Large-Scale Batch Automation & Task Management Architecture](https://uno-kim.tistory.com/478)
|
|
358
|
-
13. [[Comparison-7] #11. STT + Diarization Final: Small vs Turbo & Optimization Magic (4-Model Benchmark)](https://uno-kim.tistory.com/479)
|
|
359
|
-
14. [[Development-2] #12. Domain-Specific STT Training: Whisper Tiny Fine-Tuning on CPU](https://uno-kim.tistory.com/480)
|
|
360
|
-
15. [[Comparison-8] #13. Vanilla Model vs Custom Fine-Tuned Model (Economics/News Domain)](https://uno-kim.tistory.com/481)
|
|
361
|
-
|
|
362
|
-
---
|
|
363
|
-
|
|
364
|
-
## ⚖️ Disclaimer
|
|
365
|
-
|
|
366
|
-
> **Disclaimer:**
|
|
367
|
-
> *termux-stt is an independent open-source project developed for the Android Termux environment and is not officially affiliated with, endorsed by, or sponsored by the Termux project, OpenAI, or any other third party.*
|
|
368
|
-
>
|
|
369
|
-
> *This project is an independent open-source library engineered for Android Termux and is not officially affiliated with or endorsed by the Termux project, OpenAI, or Google LLC.*
|
|
370
|
-
|
|
371
|
-
---
|
|
372
|
-
|
|
373
|
-
## 📄 License
|
|
374
|
-
|
|
375
|
-
Released under the **MIT License**. Maintained by **AMEVA Foundation & uno-km / Eunho Kim**.
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
---
|
|
379
|
-
|
|
380
|
-
## 💖 Sponsorship & Community Backing
|
|
381
|
-
|
|
382
|
-
AMEVA is an independent open-source public good governed under the **AMEVA Open-Source Foundation (AOSF)**. All sponsorship funds are 100% publicly audited and dedicated to physical ARM64 testbeds and CI/CD GPU runners.
|
|
65
|
+
## 3. License
|
|
383
66
|
|
|
384
|
-
|
|
385
|
-
- **GitHub Sponsors**: [https://github.com/sponsors/uno-km](https://github.com/sponsors/uno-km)
|
|
386
|
-
- **Official Foundation Portal**: [https://uno-km.vercel.app/docs/foundation/sponsorship.html](https://uno-km.vercel.app/docs/foundation/sponsorship.html)
|
|
387
|
-
=======
|
|
388
|
-
Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).
|
|
389
|
-
>>>>>>> Stashed changes
|
|
67
|
+
Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).\n
|
package/index.js
CHANGED
|
@@ -21,21 +21,7 @@ function createEngine(engineName = 'whisper', options = {}) {
|
|
|
21
21
|
}
|
|
22
22
|
}
|
|
23
23
|
|
|
24
|
-
class TermuxSTT {
|
|
25
|
-
constructor(options = {}) {
|
|
26
|
-
const engineName = options.engine || 'whisper';
|
|
27
|
-
this.engine = createEngine(engineName, options);
|
|
28
|
-
}
|
|
29
|
-
transcribe(filePath, options = {}) {
|
|
30
|
-
return this.engine.transcribe(filePath, options);
|
|
31
|
-
}
|
|
32
|
-
diarize(filePath, options = {}) {
|
|
33
|
-
return this.engine.diarize(filePath, options);
|
|
34
|
-
}
|
|
35
|
-
}
|
|
36
|
-
|
|
37
24
|
module.exports = {
|
|
38
|
-
TermuxSTT,
|
|
39
25
|
createEngine,
|
|
40
26
|
Engine,
|
|
41
27
|
TranscriptResult,
|
package/lib/whisper.js
CHANGED
|
@@ -12,7 +12,15 @@ class WhisperEngine extends Engine {
|
|
|
12
12
|
super(config);
|
|
13
13
|
this.model = config.model || 'base';
|
|
14
14
|
this.lang = config.lang || 'ko';
|
|
15
|
+
this.device = config.device || 'auto';
|
|
15
16
|
this.threads = config.threads || 4;
|
|
17
|
+
|
|
18
|
+
try {
|
|
19
|
+
const avr = require('@ameva/runtime');
|
|
20
|
+
this.ctx = avr.getOrCreateContext({ device: this.device });
|
|
21
|
+
} catch (e) {
|
|
22
|
+
this.ctx = null;
|
|
23
|
+
}
|
|
16
24
|
}
|
|
17
25
|
|
|
18
26
|
async transcribe(audioPath, options = {}) {
|
|
@@ -24,6 +32,7 @@ class WhisperEngine extends Engine {
|
|
|
24
32
|
'--engine', 'whisper',
|
|
25
33
|
'--model', this.model,
|
|
26
34
|
'--lang', this.lang,
|
|
35
|
+
'--device', String(this.device),
|
|
27
36
|
'--format', 'json',
|
|
28
37
|
audioPath
|
|
29
38
|
];
|
|
@@ -54,12 +63,19 @@ class WhisperEngine extends Engine {
|
|
|
54
63
|
}
|
|
55
64
|
|
|
56
65
|
getInfo() {
|
|
57
|
-
|
|
66
|
+
const info = {
|
|
58
67
|
name: 'whisper.cpp (Node.js)',
|
|
59
68
|
model: this.model,
|
|
60
69
|
language: this.lang,
|
|
70
|
+
device: this.device,
|
|
61
71
|
threads: this.threads
|
|
62
72
|
};
|
|
73
|
+
if (this.ctx) {
|
|
74
|
+
info.backendType = this.ctx.backendType;
|
|
75
|
+
info.isGpu = this.ctx.isGpu;
|
|
76
|
+
info.deviceName = this.ctx.deviceName;
|
|
77
|
+
}
|
|
78
|
+
return info;
|
|
63
79
|
}
|
|
64
80
|
}
|
|
65
81
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "termux-stt",
|
|
3
|
-
"version": "1.1.
|
|
3
|
+
"version": "1.1.7",
|
|
4
4
|
"description": "On-device Speech-to-Text & speaker diarization framework utilizing device resources for Android Termux",
|
|
5
5
|
"main": "index.js",
|
|
6
6
|
"types": "index.d.ts",
|
|
@@ -46,9 +46,6 @@
|
|
|
46
46
|
"url": "https://github.com/uno-km/termux-stt/issues"
|
|
47
47
|
},
|
|
48
48
|
"homepage": "https://uno-km.github.io/termux-stt/",
|
|
49
|
-
"dependencies": {
|
|
50
|
-
"ameva-vulkan-runtime": ">=1.0.0"
|
|
51
|
-
},
|
|
52
49
|
"engines": {
|
|
53
50
|
"node": ">=16.0.0"
|
|
54
51
|
}
|
package/README.pypi.md
DELETED
|
@@ -1,34 +0,0 @@
|
|
|
1
|
-
# termux-stt
|
|
2
|
-
|
|
3
|
-
> **On-Device Hybrid Speech-to-Text & Speaker Diarization Engine for Android Termux**
|
|
4
|
-
> *Whisper.cpp · Vosk · Sherpa-ONNX · Non-Root ARM64 Execution · 128d Vector Diarization*
|
|
5
|
-
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
## ⚡ 5-Minute Quickstart
|
|
9
|
-
|
|
10
|
-
### Python Installation
|
|
11
|
-
|
|
12
|
-
`ash
|
|
13
|
-
# In Android Termux:
|
|
14
|
-
pkg update && pkg install -y python ffmpeg git
|
|
15
|
-
pip install termux-stt
|
|
16
|
-
`
|
|
17
|
-
|
|
18
|
-
### Python SDK Usage
|
|
19
|
-
|
|
20
|
-
`python
|
|
21
|
-
from termux_stt import STTEngine
|
|
22
|
-
|
|
23
|
-
engine = STTEngine(backend="whisper")
|
|
24
|
-
result = engine.transcribe("sample.wav")
|
|
25
|
-
print("Transcription:", result.text)
|
|
26
|
-
`
|
|
27
|
-
|
|
28
|
-
---
|
|
29
|
-
|
|
30
|
-
## 📚 Official Documentation
|
|
31
|
-
|
|
32
|
-
- **Official Web Documentation**: [https://uno-km.vercel.app/lib/stt/](https://uno-km.vercel.app/lib/stt/)
|
|
33
|
-
- **GitHub Repository**: [https://github.com/uno-km/termux-stt](https://github.com/uno-km/termux-stt)
|
|
34
|
-
- **License**: Apache-2.0
|