termux-stt 1.1.8 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,73 +1,392 @@
1
- # termux-stt
1
+ # Termux-STT
2
2
 
3
- > **Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Node.js on Android Termux.**
3
+ <div align="center">
4
+
5
+ ```
6
+ ████████╗███████╗██████╗ ███╗ ███╗██╗ ██╗██╗ ██╗ ███████╗████████╗████████╗
7
+ ╚══██╔══╝██╔════╝██╔══██╗████╗ ████║██║ ██║╚██╗██╔╝ ██╔════╝╚══██╔══╝╚══██╔══╝
8
+ ██║ █████╗ ██████╔╝██╔████╔██║██║ ██║ ╚███╔╝ █████╗███████╗ ██║ ██║
9
+ ██║ ██╔══╝ ██╔══██╗██║╚██╔╝██║██║ ██║ ██╔██╗ ╚════╝╚════██║ ██║ ██║
10
+ ██║ ███████╗██║ ██║██║ ╚═╝ ██║╚██████╔╝██╔╝ ██╗ ███████║ ██║ ██║
11
+ ╚═╝ ╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝ ╚═╝ ╚═╝
12
+ ```
13
+
14
+ **Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux**
15
+ *Dual-Engine Architecture (Python & Node.js / TypeScript) with Native Bionic ARM64 Acceleration & 0 PyTorch Dependency*
4
16
 
5
17
  <p align="center">
6
- <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/npm/v/termux-stt.svg?style=flat-square&color=cb3837&logo=npm&logoColor=white" alt="npm Version" /></a>
7
- <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/badge/npm%20Downloads-active-cb3837?style=flat-square&logo=npm&logoColor=white" alt="npm Downloads" /></a>
8
- <a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg?style=flat-square" alt="License" /></a>
9
- <img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
18
+ <a href="https://pypi.org/project/termux-stt/"><img src="https://img.shields.io/pypi/v/termux-stt.svg?style=for-the-badge&color=0088ff&logo=pypi&logoColor=white" alt="PyPI Version" /></a>
19
+ <a href="https://pypi.org/project/termux-stt/"><img src="https://img.shields.io/badge/PyPI%20Downloads-active-0088ff?style=for-the-badge&logo=pypi&logoColor=white" alt="PyPI Downloads" /></a>
20
+ <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/npm/v/termux-stt.svg?style=for-the-badge&color=cb3837&logo=npm&logoColor=white" alt="npm Version" /></a>
21
+ <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/badge/npm%20Downloads-active-cb3837?style=for-the-badge&logo=npm&logoColor=white" alt="npm Downloads" /></a>
10
22
  </p>
11
23
 
24
+ <p align="center">
25
+ <a href="https://uno-km.vercel.app/lib/stt/"><img src="https://img.shields.io/badge/Official_Docs-uno--km.vercel.app%2Flib%2Fstt-004499?style=for-the-badge&logo=vercel&logoColor=white" alt="Live Docs" /></a>
26
+ <a href="https://uno-km.vercel.app/lib/stt/demo.html"><img src="https://img.shields.io/badge/Live_Showcase-▶_Audio_Player-00f5d4?style=for-the-badge&logo=googlechrome&logoColor=0b132b" alt="Live Audio Showcase" /></a>
27
+ <a href="https://github.com/uno-km/termux-stt"><img src="https://img.shields.io/github/stars/uno-km/termux-stt?style=for-the-badge&color=gold&logo=github" alt="GitHub Stars" /></a>
28
+ <a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg?style=for-the-badge" alt="License" /></a>
29
+ </p>
30
+
31
+ <p align="center">
32
+ <img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64%2Faarch64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
33
+ <img src="https://img.shields.io/badge/Engines-whisper.cpp%20%7C%20Vosk%20%7C%20Sherpa--ONNX-38bdf8?style=flat-square" alt="Engines" />
34
+ <img src="https://img.shields.io/badge/Diarization-128d%20X--Vector%20+%20Pure%20Python-a855f7?style=flat-square" alt="Diarization" />
35
+ <img src="https://img.shields.io/badge/RAM-Under%20350MB%20(Tiny%2FBase)-10b981?style=flat-square&logo=shield&logoColor=white" alt="RAM" />
36
+ <img src="https://img.shields.io/badge/Foundation-AOSF_Tier_1-orange?style=flat-square" alt="Foundation" />
37
+ </p>
38
+
39
+ <br/>
40
+
41
+ **[Live Audio Showcase & Demo](https://uno-km.vercel.app/lib/stt/demo.html)** • **[Official Documentation Site (13 Languages)](https://uno-km.vercel.app/lib/stt/)** • **[AMEVA Foundation](https://uno-km.vercel.app/docs/foundation/)** • **[Quickstart](#1-quick-scenario-playbook)** • **[Architecture](#2-why-termux-stt-architectural-pillars)** • **[Benchmarks](#3-empirical-benchmarks-galaxy-a35--exynos-1380)**
42
+
43
+ </div>
44
+
12
45
  ---
13
46
 
14
- ## 1. Quick Start
47
+ ## AMEVA Foundation — Mobile AI Ecosystem
48
+
49
+ > **"$0 Cloud Egress, 0% External Data Leaks. Transforming every smartphone into an independent on-device AI workstation."**
50
+ > The **AMEVA Open-Source Foundation (AOSF)** builds next-generation, client-centric AI runtimes spanning on-device large models, browser automation, neural network training, and speaker diarization.
51
+
52
+ <div align="center">
15
53
 
16
- ### 1.1 Installation
54
+ | Project | Platform & Packages | Core Capability & Technology | Documentation & Demo |
55
+ | :--- | :--- | :--- | :---: |
56
+ | 🎙️ **[termux-stt](https://github.com/uno-km/termux-stt)** | [![Open Collective](https://img.shields.io/badge/Open_Collective-AOSF_Fund-004499?style=flat&logo=opencollective)](https://opencollective.com/ameva-fund) [![GitHub Sponsors](https://img.shields.io/badge/GitHub_Sponsors-uno--km-ea4aaa?style=flat&logo=githubsponsors)](https://github.com/sponsors/uno-km)<br/>[![PyPI](https://img.shields.io/pypi/v/termux-stt?color=blue&style=flat-square)](https://pypi.org/project/termux-stt/) [![npm](https://img.shields.io/npm/v/termux-stt?color=red&style=flat-square)](https://www.npmjs.com/package/termux-stt) | **Integrated On-Device STT & Pure Python 128d X-Vector Diarization** (Whisper + Vosk + Sherpa) | **[Showcase](https://uno-km.github.io/termux-stt/showcase.html)** • **[Docs](https://uno-km.vercel.app/lib/stt/)** |
57
+ | 🎨 **[termux-diffusion](https://github.com/uno-km/termux-diffusion)** | [![PyPI](https://img.shields.io/pypi/v/termux-diffusion?color=blue&style=flat-square)](https://pypi.org/project/termux-diffusion/) [![npm](https://img.shields.io/npm/v/termux-diffusion?color=red&style=flat-square)](https://www.npmjs.com/package/termux-diffusion) | **Mobile On-Device Stable Diffusion Image Generation** (bfloat16 ARM NEON acceleration) | **[Docs](https://uno-km.github.io/termux-diffusion/)** |
58
+ | 🌐 **[termux-playwright](https://github.com/uno-km/termux-playwright)** | [![PyPI](https://img.shields.io/pypi/v/termux-playwright?color=blue&style=flat-square)](https://pypi.org/project/termux-playwright/) [![npm](https://img.shields.io/npm/v/termux-playwright?color=red&style=flat-square)](https://www.npmjs.com/package/termux-playwright) | **Non-Root Native Headless Chromium Browser Automation & Scraping** | **[Docs](https://uno-km.github.io/termux-playwright/)** |
59
+ | 🧠 **[termux-train](https://github.com/uno-km/termux-train)** | [![PyPI](https://img.shields.io/pypi/v/termux-train.svg?color=blue&style=flat-square)](https://pypi.org/project/termux-train/) | **Mobile Native Autograd Neural Network Training & LoRA Fine-Tuning** | **[Docs](https://uno-km.vercel.app/lib/train/)** |
60
+ | 🖥️ **[AMEVA Workstation](https://github.com/uno-km/AMEVA-Workstation-Web)** | [![WebGPU](https://img.shields.io/badge/WebGPU-100%25_On--Device-00f5d4?style=flat-square)](https://ameva-workstation-web-core.vercel.app/) | **100% Client-Side WebGPU Multimodal Document Intelligence Workspace** | **[Live Demo](https://ameva-workstation-web-core.vercel.app/)** |
61
+ | ⚡ **[AMEVA-Forge](https://github.com/uno-km/ameva-forge)** | [![WebGPU](https://img.shields.io/badge/3D_Studio-WebGPU-purple?style=flat-square)](https://uno-km.github.io/ameva-forge/demo.html) | **Real-Time 3D Neural Studio & WebGPU Visualization Engine** | **[Live Demo](https://uno-km.github.io/ameva-forge/demo.html)** |
17
62
 
18
- #### 1-Click Automated Setup:
63
+ </div>
64
+
65
+ ---
66
+
67
+ ## 1. Installation & Provisioning
68
+
69
+ ### Method A: 1-Click Automated Installation (Recommended)
70
+ `termux-stt install` automatically provisions lightweight system dependencies (`ffmpeg`, `libbluray`, `libxml2`) and fetches the **pre-compiled Bionic ARM64 `whisper-cli` binary from GitHub Releases in <1s** (with automatic local Clang/NEON compilation fallback if offline).
71
+
72
+ #### Python SDK:
73
+ ```bash
74
+ pkg update -y && pkg install python ffmpeg git -y
75
+ pip install --upgrade termux-stt && termux-stt install
76
+ ```
77
+
78
+ #### Node.js / TypeScript:
19
79
  ```bash
20
80
  pkg update -y && pkg install nodejs-lts ffmpeg git -y
21
- npm install -g termux-stt
22
- termux-stt install
81
+ npm install -g termux-stt && termux-stt install
23
82
  ```
24
83
 
25
- #### Manual Pre-Compiled Binary Download (Direct GitHub Release):
84
+ ---
85
+
86
+ ### Method B: Manual Pre-Compiled Binary Download (Direct GitHub Release)
87
+ If you prefer to manually download the pre-compiled ARM64 native binary without building from source:
88
+
26
89
  ```bash
90
+ # 1. Download official ARM64 Bionic binary directly into ~/.local/bin
27
91
  mkdir -p ~/.local/bin
28
92
  curl -sL "https://github.com/uno-km/termux-stt/releases/download/v1.1.1/whisper-cli-arm64-android" -o ~/.local/bin/whisper-cli
29
93
  chmod 755 ~/.local/bin/whisper-cli
30
94
  ln -sf ~/.local/bin/whisper-cli ~/.local/bin/whisper-cpp
95
+
96
+ # 2. Ensure PATH includes ~/.local/bin
31
97
  export PATH=$HOME/.local/bin:$PATH
98
+
99
+ # 3. Verify installation
100
+ whisper-cli --help
101
+ ```
102
+
103
+ ---
104
+
105
+ ### Method C: Manual Source Compilation (From Scratch)
106
+ If you wish to compile `whisper.cpp` directly with Snapdragon / Cortex NEON vector optimizations:
107
+
108
+ ```bash
109
+ pkg install -y clang cmake make git ffmpeg
110
+ git clone --depth 1 https://github.com/ggerganov/whisper.cpp.git ~/tmp/whisper.cpp
111
+ cmake -B ~/tmp/whisper.cpp/build -S ~/tmp/whisper.cpp -DWHISPER_NEON=ON -DCMAKE_BUILD_TYPE=Release
112
+ cmake --build ~/tmp/whisper.cpp/build -j$(nproc)
113
+ cp ~/tmp/whisper.cpp/build/bin/whisper-cli ~/.local/bin/whisper-cli
114
+ chmod 755 ~/.local/bin/whisper-cli
32
115
  ```
33
116
 
34
- ### 1.2 Zero-Configuration CLI Demo
117
+ ---
118
+
119
+ ### Quick Verification: Out-of-the-Box Demo (JFK 1-Minute Speech)
35
120
 
36
- `termux-stt` supports instant automated demonstration without requiring manual audio preparation:
121
+ `termux-stt` provides built-in automated caching for John F. Kennedy's 1961 Inaugural Address (60.00s 16kHz Mono PCM) for instant out-of-the-box verification without requiring manual audio preparation.
37
122
 
123
+ #### Option A: Zero-Configuration Demo One-Liner (Python / npm)
38
124
  ```bash
39
- # 1. Run zero-configuration demo (automatically downloads and caches benchmark audio)
125
+ # 1. Run instant demo (automatically downloads and caches benchmark audio if missing)
40
126
  termux-stt demo
41
127
 
42
- # 2. Run demo with custom model and SRT output
128
+ # 2. Run demo with custom model and subtitle generation
43
129
  termux-stt demo --model tiny --format srt
44
130
 
45
- # 3. Transcribe custom audio
46
- termux-stt transcribe <audio_file> --engine whisper --model tiny --format srt
131
+ # 3. Transcribe demo audio directly via transcribe command
132
+ termux-stt transcribe --demo --model tiny --format text
47
133
 
48
- # 4. Transcribe with 128d Pure Python Speaker Diarization
49
- termux-stt diarize <audio_file> --speakers 2
134
+ # 4. Transcribe and separate speakers (128d X-Vector Diarization)
135
+ termux-stt diarize samples/jfk_1min.wav --speakers 2 --format text
136
+
137
+ # 5. Benchmark on-device latency & RTF
138
+ termux-stt benchmark --audio samples/jfk_1min.wav --model tiny
50
139
  ```
51
140
 
52
- ---
141
+ #### Option B: Python SDK
142
+ ```python
143
+ from termux_stt import create_engine
53
144
 
54
- ## 2. Programmatic Node.js API
145
+ # 1. Initialize Whisper Engine (auto-loads native ARM NEON binary)
146
+ engine = create_engine("whisper", model="tiny", lang="en", threads=4)
55
147
 
148
+ # 2. Transcribe JFK 60s speech
149
+ result = engine.transcribe("samples/jfk_1min.wav")
150
+
151
+ print("Transcript:\n", result.text)
152
+ print("SRT Subtitles:\n", result.to_srt())
153
+ ```
154
+
155
+ #### Option C: Node.js / TypeScript
56
156
  ```javascript
57
- const { createEngine } = require('termux-stt');
157
+ const { createEngine } = require("termux-stt");
58
158
 
59
159
  async function main() {
60
- const engine = createEngine('whisper', { model: 'tiny', lang: 'en', threads: 4 });
61
- const result = await engine.transcribe('samples/jfk_1min.wav');
62
- console.log('Transcription:', result.text);
63
- console.log('SRT Subtitles:\n', result.toSrt());
160
+ const engine = createEngine("whisper", { model: "tiny", lang: "en", threads: 4 });
161
+ const result = await engine.transcribe("samples/jfk_1min.wav");
162
+
163
+ console.log("Transcript:", result.text);
164
+ console.log("SRT Subtitles:\n", result.toSrt());
64
165
  }
65
-
66
166
  main();
67
167
  ```
68
168
 
69
169
  ---
70
170
 
71
- ## 3. License
171
+ ## 2. Advanced Scenarios
172
+
173
+ ### [Streaming] Scenario 3: Real-Time Microphone Streaming
174
+
175
+ Speak into your smartphone microphone and receive real-time transcribed text with sub-second latency:
176
+
177
+ ```python
178
+ from termux_stt import create_engine
179
+
180
+ # Use ultra-lightweight Tiny model (RTF 0.80 on Exynos 1380)
181
+ engine = create_engine("whisper", model="tiny", lang="ko")
182
+
183
+ print("🎙️ Listening... Speak into your phone microphone (Ctrl+C to stop)")
184
+ for segment in engine.stream_mic():
185
+ print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")
186
+ ```
187
+
188
+ ---
189
+
190
+ ### [Diarize] Scenario 4: Hybrid Speaker Diarization ("Who Spoke When?")
191
+
192
+ Run high-precision speaker diarization without PyTorch or CUDA:
193
+
194
+ ```python
195
+ from termux_stt import create_engine
196
+
197
+ # Hybrid Pipeline: Vosk 128d X-Vector + Whisper STT + Pure Python K-Means
198
+ engine = create_engine("hybrid", lang="ko", num_speakers=2)
199
+ result = engine.diarize("interview.wav")
200
+
201
+ for seg in result.segments:
202
+ print(f"[{seg.speaker}] ({seg.start:.1f}s - {seg.end:.1f}s): {seg.text}")
203
+ ```
204
+
205
+ *Output Example:*
206
+ ```text
207
+ [Speaker_0] (0.0s - 3.5s): Welcome to today's financial intelligence briefing.
208
+ [Speaker_1] (3.8s - 7.2s): Major equity indices closed higher following sustained foreign institutional inflows.
209
+ [Speaker_0] (7.5s - 10.1s): What is the latest outlook on the semiconductor sector?
210
+ ```
211
+
212
+ ---
213
+
214
+ ## 2. 🏛️ Why termux-stt? Architectural Pillars
215
+
216
+ ```
217
+ ┌─────────────────────────────────────────────────────────────────────────────┐
218
+ │ termux-stt Architecture │
219
+ ├────────────────────────┬────────────────────────┬───────────────────────────┤
220
+ │ User Interface Layer │ Engine Abstraction │ Output / Export Layer │
221
+ │ • Python API │ • EngineRegistry │ • JSON / SRT / VTT │
222
+ │ • Node.js API │ • create_engine() │ • RTTM (Diarization) │
223
+ │ • CLI (termux-stt) │ • ModelHub & Cache │ • Streaming Callback │
224
+ ├────────────────────────┼────────────────────────┼───────────────────────────┤
225
+ │ Core Pipeline Layer │
226
+ │ Audio Loader (7 Formats) ➔ Preprocessor (16kHz Mono) ➔ Silero-VAD Filter │
227
+ │ ➔ Multi-Engine STT (whisper.cpp / Vosk / Sherpa-ONNX) │
228
+ │ ➔ Hybrid Diarizer (128d X-Vector ➔ Pure Python Cosine/K-Means ➔ Time Align)│
229
+ ├─────────────────────────────────────────────────────────────────────────────┤
230
+ │ Platform & Process Isolation │
231
+ │ • Subprocess Isolation (Host Python never crashes on C++ Segfault) │
232
+ │ • MobileGuard (WakeLock, Doze Mode Bypass, Phantom Process Killer Shield) │
233
+ │ • Bionic ARM64 NEON & FP16 SIMD Acceleration │
234
+ └─────────────────────────────────────────────────────────────────────────────┘
235
+ ```
236
+
237
+ 1. **Subprocess Isolation**: C++ inference runs in isolated native subprocesses. Memory errors or Segfaults never kill the host Python/Node.js application.
238
+ 2. **Pure Python Clustering**: Cosine distance matrix and K-Means clustering are implemented with 0 external dependencies (`numpy` / `scikit-learn` optional, not required).
239
+ 3. **Automated Android Bionic Fixes**: Solves `sys.platform = 'linux'` spoofing, libvosk CFFI extraction, and ffmpeg audio format normalization (`16kHz / 1ch / PCM s16le`) under the hood.
240
+ 4. **Mobile Battery & CPU Guard**: Manages Android WakeLocks and background task states to prevent process termination when the screen locks.
241
+
242
+ ---
243
+
244
+ ## 3. 📊 Empirical Benchmarks (Galaxy A35 / Exynos 1380)
245
+
246
+ > *Measured on Samsung Galaxy A35 5G (Exynos 1380 4x A78 + 4x A55, 6GB RAM, Android 14 Termux).*
247
+
248
+ | Engine / Pipeline | Model | Peak RAM | RTF (Speed) | KO Accuracy | Diarization Support | Termux Rating |
249
+ | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
250
+ | **whisper.cpp** | `ggml-tiny` (39M) | **~150 MB** | **0.80** | 85% | ❌ External | ⭐⭐⭐⭐⭐ |
251
+ | **whisper.cpp** | `ggml-base` (74M) | ~250 MB | 1.20 | 88% | ❌ External | ⭐⭐⭐⭐ |
252
+ | **whisper.cpp** | `ggml-medium` (769M) | ~1.5 GB | 3.40 | **95%+** | ❌ External | ⭐⭐⭐ (Golden Acc) |
253
+ | **Vosk** | `small-ko-0.22` (42M) | **~100 MB** | **0.25** | 78% | ✅ 128d X-Vector | ⭐⭐⭐ |
254
+ | **Sherpa-ONNX** | `Zipformer` | ~300 MB | **0.42** | 86% | ✅ CAM++ | ⭐⭐⭐⭐ |
255
+ | **Pyannote.audio 3.1** | `diarization-3.1` | **> 3.5 GB** | 2.80~3.50 | N/A | ✅ Gold Standard | ❌ OOM Crashes |
256
+ | **termux-stt (Hybrid)** | `Vosk + Whisper Base`| **~350 MB** | **1.45** | **92%+** | **✅ Built-in K-Means** | **⭐⭐⭐⭐⭐ (Recommended)** |
257
+
258
+ ---
259
+
260
+ ## 4. ⚙️ Engine Comparison Matrix
261
+
262
+ | Feature | `whisper.cpp` | `Vosk` | `Sherpa-ONNX` | `Hybrid (Vosk+Whisper)` |
263
+ | :--- | :---: | :---: | :---: | :---: |
264
+ | **Primary Strength** | Highest Text Accuracy | Ultra-Low RAM & Fast | Ultra-Low Latency | **STT + Diarization Combined** |
265
+ | **Memory Footprint** | 150MB ~ 1.5GB | **< 100MB** | 300MB ~ 500MB | **~ 350MB** |
266
+ | **Real-Time Factor (RTF)** | 0.80 (Tiny) | **0.25 (Blazing)** | **0.42 (Fast)** | 1.45 (Full Pipeline) |
267
+ | **Speaker Diarization** | ❌ None | ⚠️ Basic X-Vector | ⚠️ CAM++ C++ | **✅ High-Precision Aligned** |
268
+ | **Recommended Use Case** | Quality Transcripts | Embedded / Low Spec | Live Voice Assistant | **Meetings / Interviews** |
269
+
270
+ ---
271
+
272
+ ## 5. 📚 Complete API Reference Summary
273
+
274
+ ### Python API
275
+
276
+ ```python
277
+ import termux_stt
278
+
279
+ # Create Engine with fine-grained control parameters
280
+ engine = termux_stt.create_engine(
281
+ engine="whisper", # "whisper" | "vosk" | "sherpa" | "hybrid"
282
+ model="base", # "tiny" | "base" | "small" | "medium" | "custom"
283
+ lang="ko", # ISO 639-1 language code
284
+ num_speakers=0, # 0 = disabled, 2+ = enable diarization
285
+ threads=4, # CPU threads (defaults to big cores count)
286
+ vad=True, # Enable Silero-VAD silence stripping
287
+ quantization="q5_1", # "f16" | "q8_0" | "q5_1" | "q4_0"
288
+ prompt="Financial briefing", # Initial decoding context / vocabulary
289
+ beam_size=5, # Beam search beam size
290
+ temperature=0.0 # Sampling temperature
291
+ )
292
+
293
+ # Transcribe File
294
+ result = engine.transcribe("audio.wav")
295
+ # Returns: TranscriptResult(text=str, segments=List[Segment], language=str, duration=float)
296
+
297
+ # Export Methods
298
+ result.to_json() # Structured JSON string
299
+ result.to_srt() # Standard SRT subtitle format
300
+ result.to_vtt() # WebVTT subtitle format
301
+ result.to_rttm() # NIST RTTM diarization format
302
+
303
+ # Stream Microphone
304
+ for seg in engine.stream_mic(duration=30.0):
305
+ print(f"[{seg.speaker}] {seg.text}")
306
+
307
+ # Speaker Diarization
308
+ diar_result = engine.diarize("meeting.wav", num_speakers=2)
309
+ ```
310
+
311
+ ### CLI Reference
312
+
313
+ ```bash
314
+ # General Syntax
315
+ termux-stt [COMMAND] [OPTIONS] [FILE]
316
+
317
+ # Commands
318
+ termux-stt transcribe [FILE] # Transcribe audio file (--prompt, --beam-size, --translate)
319
+ termux-stt listen # Real-time microphone transcription
320
+ termux-stt diarize [FILE] # Perform speaker diarization
321
+ termux-stt models list # List installed and available models
322
+ termux-stt models download [M] # Download specific model
323
+ termux-stt doctor # Run hardware and environment diagnostics
324
+ termux-stt benchmark # Run performance benchmark suite
325
+ ```
326
+
327
+ ---
328
+
329
+ ## 6. 🛠️ Troubleshooting & Android FAQs
330
+
331
+ ### Q1: `pip install vosk` fails with CMake or wheel error on Android
332
+ * **Cause**: Vosk does not publish official prebuilt aarch64-android wheels on PyPI.
333
+ * **Solution**: `termux-stt-install` automatically extracts `libvosk.so` from the official Android AAR and generates the CFFI bindings.
334
+
335
+ ### Q2: Whisper crashes on 44.1kHz stereo MP3/M4A files
336
+ * **Cause**: Whisper models strictly require single-channel 16,000Hz 16-bit PCM WAV.
337
+ * **Solution**: `termux-stt` automatically runs `ffmpeg` normalization on any audio format (`mp3`, `m4a`, `flac`, `ogg`, `opus`, `webm`).
338
+
339
+ ### Q3: Process killed after 10 minutes in background
340
+ * **Cause**: Android Phantom Process Killer terminates background tasks.
341
+ * **Solution**: Enable Termux WakeLock (`termux-wake-lock`) and disable battery optimization for Termux in Android Settings.
342
+
343
+ ---
344
+
345
+ ## 7. 🔍 15-Part Empirical Research Blog Series
346
+
347
+ This framework is built upon the exhaustive 15-part research series published on [Eunho Kim's Technical Engineering Blog](https://uno-kim.tistory.com/):
348
+
349
+ 1. [[Whisper.cpp] #1. Edge Agent AI: Whisper.cpp Speech Processing (Base vs Tiny)](https://uno-kim.tistory.com/467)
350
+ 2. [[Audio Extraction] #2. Extracting Specific Audio Segments on Android](https://uno-kim.tistory.com/468)
351
+ 3. [[Whisper.cpp] #3. Korean STT Conversion Comparison (4 Models)](https://uno-kim.tistory.com/469)
352
+ 4. [[Comparison-1] #4. STT + Speaker Diarization: 3 Lightweight Engines + Pyannote](https://uno-kim.tistory.com/472)
353
+ 5. [[Comparison-2] #5. Sherpa-ONNX Execution, Diarization, and Troubleshooting](https://uno-kim.tistory.com/473)
354
+ 6. [[Comparison-3] #6. Speaker Diarization using Pyannote Model](https://uno-kim.tistory.com/471)
355
+ 7. [[Pyannote] Troubleshooting Diarization in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/470)
356
+ 8. [[Comparison-4] #7. Vosk Execution, Speaker Diarization, and Troubleshooting](https://uno-kim.tistory.com/475)
357
+ 9. [[Vosk] Troubleshooting Vosk in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/474)
358
+ 10. [[Comparison-5] #8. Vosk / Pyannote / Sherpa-ONNX / Whisper.cpp Comprehensive Comparison](https://uno-kim.tistory.com/476)
359
+ 11. [[Comparison-6] #9. Vosk + Whisper.cpp Hybrid Pipeline & X-Vector Diarization](https://uno-kim.tistory.com/477)
360
+ 12. [[Development-1] #10. Large-Scale Batch Automation & Task Management Architecture](https://uno-kim.tistory.com/478)
361
+ 13. [[Comparison-7] #11. STT + Diarization Final: Small vs Turbo & Optimization Magic (4-Model Benchmark)](https://uno-kim.tistory.com/479)
362
+ 14. [[Development-2] #12. Domain-Specific STT Training: Whisper Tiny Fine-Tuning on CPU](https://uno-kim.tistory.com/480)
363
+ 15. [[Comparison-8] #13. Vanilla Model vs Custom Fine-Tuned Model (Economics/News Domain)](https://uno-kim.tistory.com/481)
364
+
365
+ ---
366
+
367
+ ## ⚖️ Disclaimer
368
+
369
+ > **Disclaimer:**
370
+ > *termux-stt is an independent open-source project developed for the Android Termux environment and is not officially affiliated with, endorsed by, or sponsored by the Termux project, OpenAI, or any other third party.*
371
+ >
372
+ > *This project is an independent open-source library engineered for Android Termux and is not officially affiliated with or endorsed by the Termux project, OpenAI, or Google LLC.*
373
+
374
+ ---
375
+
376
+ ## 📄 License
377
+
378
+ Released under the **MIT License**. Maintained by **AMEVA Foundation & uno-km / Eunho Kim**.
379
+
380
+
381
+ ---
382
+
383
+ ## 💖 Sponsorship & Community Backing
384
+
385
+ AMEVA is an independent open-source public good governed under the **AMEVA Open-Source Foundation (AOSF)**. All sponsorship funds are 100% publicly audited and dedicated to physical ARM64 testbeds and CI/CD GPU runners.
72
386
 
73
- Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).\n
387
+ - **Open Collective (Non-Profit 501(c)(6))**: [https://opencollective.com/ameva-fund](https://opencollective.com/ameva-fund)
388
+ - **GitHub Sponsors**: [https://github.com/sponsors/uno-km](https://github.com/sponsors/uno-km)
389
+ - **Official Foundation Portal**: [https://uno-km.vercel.app/docs/foundation/sponsorship.html](https://uno-km.vercel.app/docs/foundation/sponsorship.html)
390
+ =======
391
+ Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).
392
+ >>>>>>> Stashed changes
package/README.pypi.md ADDED
@@ -0,0 +1,42 @@
1
+ # termux-stt
2
+
3
+ > **On-Device Hybrid Speech-to-Text & Speaker Diarization Engine for Android Termux**
4
+ > *Whisper.cpp · Vosk · Sherpa-ONNX · Non-Root ARM64 Execution · 128d Vector Diarization*
5
+
6
+ ---
7
+
8
+ ## 5-Minute Quickstart
9
+
10
+ ### Installation
11
+
12
+ ```bash
13
+ # In Android Termux:
14
+ pkg update && pkg install -y python ffmpeg git
15
+ pip install --upgrade termux-stt
16
+ termux-stt install
17
+ ```
18
+
19
+ ### Instant CLI Demo
20
+
21
+ ```bash
22
+ # Automatic benchmark audio download & transcription
23
+ termux-stt demo
24
+ ```
25
+
26
+ ### Python SDK Usage
27
+
28
+ ```python
29
+ from termux_stt import create_engine
30
+
31
+ engine = create_engine("whisper", model="tiny", lang="en")
32
+ result = engine.transcribe("sample.wav")
33
+ print("Transcription:", result.text)
34
+ ```
35
+
36
+ ---
37
+
38
+ ## 📚 Official Documentation
39
+
40
+ - **Official Web Documentation**: [https://uno-km.vercel.app/lib/stt/](https://uno-km.vercel.app/lib/stt/)
41
+ - **GitHub Repository**: [https://github.com/uno-km/termux-stt](https://github.com/uno-km/termux-stt)
42
+ - **License**: Apache-2.0
package/bin/termux-stt.js CHANGED
@@ -9,10 +9,13 @@ const { spawn, spawnSync } = require('child_process');
9
9
  const args = process.argv.slice(2);
10
10
 
11
11
  function getPythonExecutable() {
12
- const candidates = ['python3', 'python'];
12
+ if (process.env.PYTHON) return process.env.PYTHON;
13
+ const candidates = process.platform === 'win32'
14
+ ? ['py', 'python3', 'python']
15
+ : ['python3', 'python'];
13
16
  for (const cmd of candidates) {
14
17
  try {
15
- const res = spawnSync(cmd, ['--version'], { stdio: 'ignore' });
18
+ const res = spawnSync(cmd, ['-c', 'import sys; sys.exit(0)'], { stdio: 'ignore' });
16
19
  if (res.status === 0) return cmd;
17
20
  } catch (_) {}
18
21
  }
package/index.d.ts CHANGED
@@ -1,6 +1,7 @@
1
- export interface EngineOptions {
1
+ export interface EngineOptions {
2
2
  model?: string;
3
3
  lang?: string;
4
+ device?: 'auto' | 'vulkan' | 'gpu' | 'cpu';
4
5
  threads?: number;
5
6
  vad?: boolean;
6
7
  vadThreshold?: number;
package/index.js CHANGED
@@ -21,7 +21,21 @@ function createEngine(engineName = 'whisper', options = {}) {
21
21
  }
22
22
  }
23
23
 
24
+ class TermuxSTT {
25
+ constructor(options = {}) {
26
+ const engineName = options.engine || 'whisper';
27
+ this.engine = createEngine(engineName, options);
28
+ }
29
+ transcribe(filePath, options = {}) {
30
+ return this.engine.transcribe(filePath, options);
31
+ }
32
+ diarize(filePath, options = {}) {
33
+ return this.engine.diarize(filePath, options);
34
+ }
35
+ }
36
+
24
37
  module.exports = {
38
+ TermuxSTT,
25
39
  createEngine,
26
40
  Engine,
27
41
  TranscriptResult,
package/lib/whisper.js CHANGED
@@ -12,32 +12,35 @@ class WhisperEngine extends Engine {
12
12
  super(config);
13
13
  this.model = config.model || 'base';
14
14
  this.lang = config.lang || 'ko';
15
+ this.threads = config.threads || null;
15
16
  this.device = config.device || 'auto';
16
- this.threads = config.threads || 4;
17
-
18
- try {
19
- const avr = require('@ameva/runtime');
20
- this.ctx = avr.getOrCreateContext({ device: this.device });
21
- } catch (e) {
22
- this.ctx = null;
23
- }
24
17
  }
25
18
 
26
19
  async transcribe(audioPath, options = {}) {
27
20
  return new Promise((resolve, reject) => {
28
- // Use termux-stt python CLI or direct whisper.cpp
21
+ const device = options.device || this.device || 'auto';
22
+ const model = options.model || this.model || 'base';
23
+ const lang = options.lang || this.lang || 'ko';
24
+ const threads = options.threads || this.threads;
25
+
29
26
  const args = [
30
27
  '-m', 'termux_stt.cli.main',
31
28
  'transcribe',
32
29
  '--engine', 'whisper',
33
- '--model', this.model,
34
- '--lang', this.lang,
35
- '--device', String(this.device),
30
+ '--model', model,
31
+ '--lang', lang,
32
+ '--device', device,
36
33
  '--format', 'json',
37
- audioPath
38
34
  ];
39
35
 
40
- const proc = spawn('python', args, { env: process.env });
36
+ if (threads) {
37
+ args.push('--threads', String(threads));
38
+ }
39
+
40
+ args.push(audioPath);
41
+
42
+ const pythonExe = process.env.PYTHON || (process.platform === 'win32' ? 'python' : 'python3');
43
+ const proc = spawn(pythonExe, args, { env: process.env });
41
44
  let stdout = '';
42
45
  let stderr = '';
43
46
 
@@ -53,29 +56,23 @@ class WhisperEngine extends Engine {
53
56
  const segments = (parsed.segments || []).map(
54
57
  s => new Segment(s.start, s.end, s.text, s.speaker, s.confidence)
55
58
  );
56
- resolve(new TranscriptResult(parsed.text, segments, parsed.language || this.lang, parsed.duration));
59
+ resolve(new TranscriptResult(parsed.text, segments, parsed.language || lang, parsed.duration));
57
60
  } catch (e) {
58
61
  // Fallback if stdout was plain text
59
- resolve(new TranscriptResult(stdout.trim(), [new Segment(0, 0, stdout.trim())], this.lang));
62
+ resolve(new TranscriptResult(stdout.trim(), [new Segment(0, 0, stdout.trim())], lang));
60
63
  }
61
64
  });
62
65
  });
63
66
  }
64
67
 
65
68
  getInfo() {
66
- const info = {
69
+ return {
67
70
  name: 'whisper.cpp (Node.js)',
68
71
  model: this.model,
69
72
  language: this.lang,
73
+ threads: this.threads,
70
74
  device: this.device,
71
- threads: this.threads
72
75
  };
73
- if (this.ctx) {
74
- info.backendType = this.ctx.backendType;
75
- info.isGpu = this.ctx.isGpu;
76
- info.deviceName = this.ctx.deviceName;
77
- }
78
- return info;
79
76
  }
80
77
  }
81
78
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "termux-stt",
3
- "version": "1.1.8",
3
+ "version": "1.2.0",
4
4
  "description": "On-device Speech-to-Text & speaker diarization framework utilizing device resources for Android Termux",
5
5
  "main": "index.js",
6
6
  "types": "index.d.ts",
@@ -45,6 +45,9 @@
45
45
  "url": "https://github.com/uno-km/termux-stt/issues"
46
46
  },
47
47
  "homepage": "https://uno-km.github.io/termux-stt/",
48
+ "dependencies": {
49
+ "@ameva/runtime": ">=2.0.0"
50
+ },
48
51
  "engines": {
49
52
  "node": ">=16.0.0"
50
53
  }