termux-stt 1.0.0 → 1.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +86 -416
  2. package/package.json +50 -50
package/README.md CHANGED
@@ -1,416 +1,86 @@
1
- # Termux-STT
2
-
3
- <div align="center">
4
-
5
- ```
6
- ████████╗███████╗██████╗ ███╗ ███╗██╗ ██╗██╗ ██╗ ███████╗████████╗████████╗
7
- ╚══██╔══╝██╔════╝██╔══██╗████╗ ████║██║ ██║╚██╗██╔╝ ██╔════╝╚══██╔══╝╚══██╔══╝
8
- ██║ █████╗ ██████╔╝██╔████╔██║██║ ██║ ╚███╔╝ █████╗███████╗ ██║ ██║
9
- ██║ ██╔══╝ ██╔══██╗██║╚██╔╝██║██║ ██║ ██╔██╗ ╚════╝╚════██║ ██║ ██║
10
- ██║ ███████╗██║ ██║██║ ╚═╝ ██║╚██████╔╝██╔╝ ██╗ ███████║ ██║ ██║
11
- ╚═╝ ╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝ ╚═╝ ╚═╝
12
- ```
13
-
14
- **Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux**
15
- *Dual-Engine Architecture (Python & Node.js / TypeScript) with Native Bionic ARM64 Acceleration & 0 PyTorch Dependency*
16
-
17
- <p align="center">
18
- <a href="https://pypi.org/project/termux-stt/"><img src="https://img.shields.io/pypi/v/termux-stt.svg?color=blue&style=for-the-badge&logo=pypi&logoColor=white" alt="PyPI Version" /></a>
19
- <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/npm/v/termux-stt.svg?color=red&style=for-the-badge&logo=npm&logoColor=white" alt="npm Version" /></a>
20
- <a href="https://uno-km.github.io/termux-stt/showcase.html"><img src="https://img.shields.io/badge/Live_Showcase-▶_Audio_Player-00f5d4?style=for-the-badge&logo=googlechrome&logoColor=0b132b" alt="Live Audio Showcase" /></a>
21
- <a href="https://uno-km.github.io/termux-stt/"><img src="https://img.shields.io/badge/Docs-uno--km.github.io-004499?style=for-the-badge&logo=googlechrome&logoColor=white" alt="Live Docs" /></a>
22
- <a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-MIT-green.svg?style=for-the-badge" alt="License" /></a>
23
- </p>
24
-
25
- <p align="center">
26
- <img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64%2Faarch64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
27
- <img src="https://img.shields.io/badge/Engines-whisper.cpp%20%7C%20Vosk%20%7C%20Sherpa--ONNX-38bdf8?style=flat-square" alt="Engines" />
28
- <img src="https://img.shields.io/badge/Diarization-128d%20X--Vector%20+%20Pure%20Python-a855f7?style=flat-square" alt="Diarization" />
29
- <img src="https://img.shields.io/badge/RAM-Under%20350MB%20(Tiny%2FBase)-10b981?style=flat-square&logo=shield&logoColor=white" alt="RAM" />
30
- <img src="https://img.shields.io/badge/Python-3.8%20%7C%203.9%20%7C%203.10%20%7C%203.11%20%7C%203.12-f59e0b?style=flat-square&logo=python&logoColor=white" alt="Python" />
31
- <img src="https://img.shields.io/badge/Node.js-16%20%7C%2018%20%7C%2020%20%7C%2022-3178c6?style=flat-square&logo=nodedotjs&logoColor=white" alt="Node" />
32
- </p>
33
-
34
- <br/>
35
-
36
- **[🎧 Live Audio Showcase & Demo](https://uno-km.github.io/termux-stt/showcase.html)** • **[📖 Official Documentation Site](https://uno-km.github.io/termux-stt/)** • **[⚡ Quickstart](#1-quick-scenario-playbook)** • **[🏛️ Architecture](#2-why-termux-stt-architectural-pillars)** • **[📊 Benchmarks](#3-empirical-benchmarks-galaxy-a35--exynos-1380)** • **[🔍 15-Part Blog Series](#7-15-part-empirical-research-blog-series)**
37
-
38
- </div>
39
-
40
- ---
41
-
42
- ## 🎙️ 단 3줄로 끝내는 안드로이드 Termux 온디바이스 음성인식
43
-
44
- ```python
45
- from termux_stt import create_engine
46
-
47
- # 1. 엔진 초기화 (최초 1회 모델 자동 다운로드 및 캐싱)
48
- engine = create_engine("whisper", model="base", lang="en")
49
-
50
- # 2. 오디오 전사 및 자막 생성 (WAV, MP3, M4A, FLAC, OGG 자동 16kHz 변환)
51
- result = engine.transcribe("speech.mp3")
52
-
53
- print(result.text) # 전체 텍스트
54
- print(result.to_srt()) # 표준 SRT 자막
55
- ```
56
-
57
- ```bash
58
- # 또는 터미널에서 1줄 CLI 실행
59
- termux-stt transcribe --engine whisper --model base speech.mp3
60
- ```
61
-
62
- ---
63
-
64
- ## 🎧 Live Audio Showcase & On-Device Transcription Proof
65
-
66
- > **[▶ 웹 브라우저에서 실시간 음성 및 동기화 자막 체험하기 (Live Audio Showcase)](https://uno-km.github.io/termux-stt/showcase.html)**
67
-
68
- ### 1. 실측 오디오 스펙 & 전사 타임라인
69
-
70
- * **입력 오디오**: `continuous_speech.wav` (37.91초, 16000Hz Mono PCM)
71
- * **추론 엔진**: `whisper.cpp Base` (On-Device Local CPU)
72
- * **처리 시간**: **32.79초** (RTF: **0.865x**, 실시간보다 빠름)
73
- * **문장 반복률**: **0%** (모든 발화 구간이 각기 다른 내용으로 고유하게 전사됨)
74
-
75
- | No. | 타임스탬프 (시작 → 종료) | 전사된 문장 (Transcribed Text) |
76
- | :---: | :---: | :--- |
77
- | **01** | `00:00.00 → 00:09.36` | *"And so my fellow Americans, ask not what your country can do for you, ask what you can"* |
78
- | **02** | `00:09.36 → 00:11.60` | *"do for your country."* |
79
- | **03** | `00:11.60 → 00:16.18` | *He hoped there would be stew for dinner, turnips and carrots and bruised potatoes and* |
80
- | **04** | `00:16.18 → 00:22.00` | *fat mutton pieces to be ladled out in thick, peppered flour-fatten sauce.* |
81
- | **05** | `00:22.00 → 00:25.36` | *Stuff it into you, his belly counseled him.* |
82
- | **06** | `00:25.36 → 00:29.88` | *After early nightfall, the yellow lamps would light up here and there, the squalid quarter* |
83
- | **07** | `00:29.88 → 00:37.14` | *of the brothels.* |
84
-
85
- ### 2. 자동 생성된 SRT 자막 파일
86
-
87
- ```srt
88
- 1
89
- 00:00:00,000 --> 00:00:09,360
90
- "And so my fellow Americans, ask not what your country can do for you, ask what you can
91
-
92
- 2
93
- 00:00:09,360 --> 00:00:11,600
94
- do for your country."
95
-
96
- 3
97
- 00:00:11,600 --> 00:00:16,180
98
- He hoped there would be stew for dinner, turnips and carrots and bruised potatoes and
99
-
100
- 4
101
- 00:00:16,180 --> 00:00:22,000
102
- fat mutton pieces to be ladled out in thick, peppered flour-fatten sauce.
103
-
104
- 5
105
- 00:00:22,000 --> 00:00:25,360
106
- Stuff it into you, his belly counseled him.
107
-
108
- 6
109
- 00:00:25,360 --> 00:00:29,880
110
- After early nightfall, the yellow lamps would light up here and there, the squalid quarter
111
-
112
- 7
113
- 00:00:29,880 --> 00:00:37,140
114
- of the brothels.
115
- ```
116
-
117
- ---
118
-
119
- ## 💡 What is termux-stt?
120
-
121
- `termux-stt` is an all-in-one, production-ready speech-to-text and speaker diarization framework engineered natively for **Android Termux (ARM64 / aarch64)**.
122
-
123
- Standard mobile STT setups force developers to endure 30+ minutes of manual CMake builds, broken PyPI wheels on Android Bionic, 2GB+ PyTorch binaries that trigger Android Low Memory Killer (OOM), and broken platform guards.
124
-
125
- **`termux-stt` eliminates all friction with a 3-line unified API:**
126
- - **Zero-PyTorch Dependency**: Replaces heavy ML frameworks with C++ binary subprocess isolation and Pure Python clustering math.
127
- - **Multi-Engine Unification**: Run `whisper.cpp`, `Vosk`, or `Sherpa-ONNX` via the exact same `create_engine()` interface.
128
- - **Built-in Hybrid Diarization**: Combines Vosk 128d X-Vector voice fingerprints with Whisper STT under 1.5 GB RAM.
129
- - **Empirically Proven**: Engineered from 15 comprehensive benchmarks on Samsung Galaxy A35 (Exynos 1380, 6GB RAM).
130
-
131
- ---
132
-
133
- ## 1. Quick Scenario Playbook
134
-
135
- ### [Install] Scenario 1: Clean Install (Fresh Setup on Android Termux)
136
-
137
- #### [Python] Python (`pip`):
138
- ```bash
139
- # 1. Grant Storage & Microphone Permissions in Termux
140
- termux-setup-storage
141
-
142
- # 2. Install Dependencies & Provision Native Engines
143
- pkg update -y && pkg install python clang make cmake git ffmpeg termux-api -y
144
- pip install termux-stt && termux-stt-install
145
- ```
146
-
147
- #### [Node.js] Node.js / TypeScript (`npm`):
148
- ```bash
149
- # 1. Grant Storage & Microphone Permissions
150
- termux-setup-storage
151
-
152
- # 2. Install Dependencies & Provision Native Engines
153
- pkg update -y && pkg install nodejs-lts clang make cmake git ffmpeg termux-api -y
154
- npm install -g termux-stt && npx termux-stt install
155
- ```
156
-
157
- ---
158
-
159
- ### [Instant] Scenario 2: Instant Transcription (Ready to Run)
160
-
161
- #### Option A: One-Line CLI
162
- ```bash
163
- # Transcribe audio file with default Whisper engine (Korean)
164
- termux-stt transcribe meeting.wav
165
-
166
- # Export directly to Subtitles (SRT or VTT)
167
- termux-stt transcribe --format srt meeting.wav > subtitles.srt
168
-
169
- # Use ultra-fast Vosk engine
170
- termux-stt transcribe --engine vosk --model small-ko voice_memo.wav
171
- ```
172
-
173
- #### Option B: Python SDK Integration
174
- ```python
175
- from termux_stt import create_engine
176
-
177
- # 1. Initialize Engine (auto-downloads model on first call)
178
- engine = create_engine("whisper", model="base", lang="ko")
179
-
180
- # 2. Transcribe Audio
181
- result = engine.transcribe("meeting.wav")
182
-
183
- print("Transcript:", result.text)
184
- print("Detected Language:", result.language)
185
- print("Duration:", f"{result.duration:.2f}s")
186
- ```
187
-
188
- #### Option C: Node.js / TypeScript Integration
189
- ```javascript
190
- const { createEngine } = require("termux-stt");
191
-
192
- async function main() {
193
- const engine = createEngine("whisper", { model: "base", lang: "ko" });
194
- const result = await engine.transcribe("meeting.wav");
195
-
196
- console.log("Transcript:", result.text);
197
- console.log("SRT Subtitles:\n", result.toSrt());
198
- }
199
- main();
200
- ```
201
-
202
- ---
203
-
204
- ### [Streaming] Scenario 3: Real-Time Microphone Streaming
205
-
206
- Speak into your smartphone microphone and receive real-time transcribed text with sub-second latency:
207
-
208
- ```python
209
- from termux_stt import create_engine
210
-
211
- # Use ultra-lightweight Tiny model (RTF 0.80 on Exynos 1380)
212
- engine = create_engine("whisper", model="tiny", lang="ko")
213
-
214
- print("🎙️ Listening... Speak into your phone microphone (Ctrl+C to stop)")
215
- for segment in engine.stream_mic():
216
- print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")
217
- ```
218
-
219
- ---
220
-
221
- ### [Diarize] Scenario 4: Hybrid Speaker Diarization ("Who Spoke When?")
222
-
223
- Run high-precision speaker diarization without PyTorch or CUDA:
224
-
225
- ```python
226
- from termux_stt import create_engine
227
-
228
- # Hybrid Pipeline: Vosk 128d X-Vector + Whisper STT + Pure Python K-Means
229
- engine = create_engine("hybrid", lang="ko", num_speakers=2)
230
- result = engine.diarize("interview.wav")
231
-
232
- for seg in result.segments:
233
- print(f"[{seg.speaker}] ({seg.start:.1f}s - {seg.end:.1f}s): {seg.text}")
234
- ```
235
-
236
- *Output Example:*
237
- ```text
238
- [Speaker_0] (0.0s - 3.5s): 오늘 경제 브리핑을 시작하겠습니다.
239
- [Speaker_1] (3.8s - 7.2s): 네, 오늘 코스피 지수가 외국인 순매수로 상승 마감했습니다.
240
- [Speaker_0] (7.5s - 10.1s): 반도체 섹터 동향은 어떤가요?
241
- ```
242
-
243
- ---
244
-
245
- ## 2. 🏛️ Why termux-stt? Architectural Pillars
246
-
247
- ```
248
- ┌─────────────────────────────────────────────────────────────────────────────┐
249
- │ termux-stt Architecture │
250
- ├────────────────────────┬────────────────────────┬───────────────────────────┤
251
- │ User Interface Layer │ Engine Abstraction │ Output / Export Layer │
252
- │ • Python API │ • EngineRegistry │ • JSON / SRT / VTT │
253
- │ • Node.js API │ • create_engine() │ • RTTM (Diarization) │
254
- │ • CLI (termux-stt) │ • ModelHub & Cache │ • Streaming Callback │
255
- ├────────────────────────┼────────────────────────┼───────────────────────────┤
256
- │ Core Pipeline Layer │
257
- │ Audio Loader (7 Formats) ➔ Preprocessor (16kHz Mono) ➔ Silero-VAD Filter │
258
- │ ➔ Multi-Engine STT (whisper.cpp / Vosk / Sherpa-ONNX) │
259
- │ ➔ Hybrid Diarizer (128d X-Vector ➔ Pure Python Cosine/K-Means ➔ Time Align)│
260
- ├─────────────────────────────────────────────────────────────────────────────┤
261
- │ Platform & Process Isolation │
262
- │ • Subprocess Isolation (Host Python never crashes on C++ Segfault) │
263
- │ • MobileGuard (WakeLock, Doze Mode Bypass, Phantom Process Killer Shield) │
264
- │ • Bionic ARM64 NEON & FP16 SIMD Acceleration │
265
- └─────────────────────────────────────────────────────────────────────────────┘
266
- ```
267
-
268
- 1. **Subprocess Isolation**: C++ inference runs in isolated native subprocesses. Memory errors or Segfaults never kill the host Python/Node.js application.
269
- 2. **Pure Python Clustering**: Cosine distance matrix and K-Means clustering are implemented with 0 external dependencies (`numpy` / `scikit-learn` optional, not required).
270
- 3. **Automated Android Bionic Fixes**: Solves `sys.platform = 'linux'` spoofing, libvosk CFFI extraction, and ffmpeg audio format normalization (`16kHz / 1ch / PCM s16le`) under the hood.
271
- 4. **Mobile Battery & CPU Guard**: Manages Android WakeLocks and background task states to prevent process termination when the screen locks.
272
-
273
- ---
274
-
275
- ## 3. 📊 Empirical Benchmarks (Galaxy A35 / Exynos 1380)
276
-
277
- > *Measured on Samsung Galaxy A35 5G (Exynos 1380 4x A78 + 4x A55, 6GB RAM, Android 14 Termux).*
278
-
279
- | Engine / Pipeline | Model | Peak RAM | RTF (Speed) | KO Accuracy | Diarization Support | Termux Rating |
280
- | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
281
- | **whisper.cpp** | `ggml-tiny` (39M) | **~150 MB** | **0.80** | 85% | ❌ External | ⭐⭐⭐⭐⭐ |
282
- | **whisper.cpp** | `ggml-base` (74M) | ~250 MB | 1.20 | 88% | ❌ External | ⭐⭐⭐⭐ |
283
- | **whisper.cpp** | `ggml-medium` (769M) | ~1.5 GB | 3.40 | **95%+** | ❌ External | ⭐⭐⭐ (Golden Acc) |
284
- | **Vosk** | `small-ko-0.22` (42M) | **~100 MB** | **0.25** | 78% | ✅ 128d X-Vector | ⭐⭐⭐ |
285
- | **Sherpa-ONNX** | `Zipformer` | ~300 MB | **0.42** | 86% | ✅ CAM++ | ⭐⭐⭐⭐ |
286
- | **Pyannote.audio 3.1** | `diarization-3.1` | **> 3.5 GB** | 2.80~3.50 | N/A | ✅ Gold Standard | ❌ OOM Crashes |
287
- | **termux-stt (Hybrid)** | `Vosk + Whisper Base`| **~350 MB** | **1.45** | **92%+** | **✅ Built-in K-Means** | **⭐⭐⭐⭐⭐ (Recommended)** |
288
-
289
- ---
290
-
291
- ## 4. ⚙️ Engine Comparison Matrix
292
-
293
- | Feature | `whisper.cpp` | `Vosk` | `Sherpa-ONNX` | `Hybrid (Vosk+Whisper)` |
294
- | :--- | :---: | :---: | :---: | :---: |
295
- | **Primary Strength** | Highest Text Accuracy | Ultra-Low RAM & Fast | Ultra-Low Latency | **STT + Diarization Combined** |
296
- | **Memory Footprint** | 150MB ~ 1.5GB | **< 100MB** | 300MB ~ 500MB | **~ 350MB** |
297
- | **Real-Time Factor (RTF)** | 0.80 (Tiny) | **0.25 (Blazing)** | **0.42 (Fast)** | 1.45 (Full Pipeline) |
298
- | **Speaker Diarization** | ❌ None | ⚠️ Basic X-Vector | ⚠️ CAM++ C++ | **✅ High-Precision Aligned** |
299
- | **Recommended Use Case** | Quality Transcripts | Embedded / Low Spec | Live Voice Assistant | **Meetings / Interviews** |
300
-
301
- ---
302
-
303
- ## 5. 📚 Complete API Reference Summary
304
-
305
- ### Python API
306
-
307
- ```python
308
- import termux_stt
309
-
310
- # Create Engine
311
- engine = termux_stt.create_engine(
312
- engine="whisper", # "whisper" | "vosk" | "sherpa" | "hybrid"
313
- model="base", # "tiny" | "base" | "small" | "medium" | "custom"
314
- lang="ko", # ISO 639-1 language code
315
- num_speakers=0, # 0 = disabled, 2+ = enable diarization
316
- threads=4, # CPU threads (defaults to big cores count)
317
- vad=True, # Enable Silero-VAD silence stripping
318
- quantization="q5_1" # "f16" | "q8_0" | "q5_1" | "q4_0"
319
- )
320
-
321
- # Transcribe File
322
- result = engine.transcribe("audio.wav")
323
- # Returns: TranscriptResult(text=str, segments=List[Segment], language=str, duration=float)
324
-
325
- # Export Methods
326
- result.to_json() # Structured JSON string
327
- result.to_srt() # Standard SRT subtitle format
328
- result.to_vtt() # WebVTT subtitle format
329
- result.to_rttm() # NIST RTTM diarization format
330
-
331
- # Stream Microphone
332
- for seg in engine.stream_mic(duration=30.0):
333
- print(f"[{seg.speaker}] {seg.text}")
334
-
335
- # Speaker Diarization
336
- diar_result = engine.diarize("meeting.wav", num_speakers=2)
337
- ```
338
-
339
- ### CLI Reference
340
-
341
- ```bash
342
- # General Syntax
343
- termux-stt [COMMAND] [OPTIONS] [FILE]
344
-
345
- # Commands
346
- termux-stt transcribe [FILE] # Transcribe audio file
347
- termux-stt listen # Real-time microphone transcription
348
- termux-stt diarize [FILE] # Perform speaker diarization
349
- termux-stt models list # List installed and available models
350
- termux-stt models download [M] # Download specific model
351
- termux-stt doctor # Run hardware and environment diagnostics
352
- termux-stt benchmark # Run performance benchmark suite
353
- ```
354
-
355
- ---
356
-
357
- ## 6. 🛠️ Troubleshooting & Android FAQs
358
-
359
- ### Q1: `pip install vosk` fails with CMake or wheel error on Android
360
- * **Cause**: Vosk does not publish official prebuilt aarch64-android wheels on PyPI.
361
- * **Solution**: `termux-stt-install` automatically extracts `libvosk.so` from the official Android AAR and generates the CFFI bindings.
362
-
363
- ### Q2: Whisper crashes on 44.1kHz stereo MP3/M4A files
364
- * **Cause**: Whisper models strictly require single-channel 16,000Hz 16-bit PCM WAV.
365
- * **Solution**: `termux-stt` automatically runs `ffmpeg` normalization on any audio format (`mp3`, `m4a`, `flac`, `ogg`, `opus`, `webm`).
366
-
367
- ### Q3: Process killed after 10 minutes in background
368
- * **Cause**: Android Phantom Process Killer terminates background tasks.
369
- * **Solution**: Enable Termux WakeLock (`termux-wake-lock`) and disable battery optimization for Termux in Android Settings.
370
-
371
- ---
372
-
373
- ## 7. 🔍 15-Part Empirical Research Blog Series
374
-
375
- This framework is built upon the exhaustive 15-part research series published on [Eunho Kim's Technical Blog (우노킴 티스토리)](https://uno-kim.tistory.com/):
376
-
377
- 1. [[Whisper.cpp] #1. Edge Agent AI: Whisper.cpp Speech Processing (Base vs Tiny)](https://uno-kim.tistory.com/467)
378
- 2. [[Audio Extraction] #2. Extracting Specific Audio Segments on Android](https://uno-kim.tistory.com/468)
379
- 3. [[Whisper.cpp] #3. Korean STT Conversion Comparison (4 Models)](https://uno-kim.tistory.com/469)
380
- 4. [[Comparison-1] #4. STT + Speaker Diarization: 3 Lightweight Engines + Pyannote](https://uno-kim.tistory.com/472)
381
- 5. [[Comparison-2] #5. Sherpa-ONNX Execution, Diarization, and Troubleshooting](https://uno-kim.tistory.com/473)
382
- 6. [[Comparison-3] #6. Speaker Diarization using Pyannote Model](https://uno-kim.tistory.com/471)
383
- 7. [[Pyannote] Troubleshooting Diarization in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/470)
384
- 8. [[Comparison-4] #7. Vosk Execution, Speaker Diarization, and Troubleshooting](https://uno-kim.tistory.com/475)
385
- 9. [[Vosk] Troubleshooting Vosk in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/474)
386
- 10. [[Comparison-5] #8. Vosk / Pyannote / Sherpa-ONNX / Whisper.cpp Comprehensive Comparison](https://uno-kim.tistory.com/476)
387
- 11. [[Comparison-6] #9. Vosk + Whisper.cpp Hybrid Pipeline & X-Vector Diarization](https://uno-kim.tistory.com/477)
388
- 12. [[Development-1] #10. Large-Scale Batch Automation & Task Management Architecture](https://uno-kim.tistory.com/478)
389
- 13. [[Comparison-7] #11. STT + Diarization Final: Small vs Turbo & Optimization Magic (4-Model Benchmark)](https://uno-kim.tistory.com/479)
390
- 14. [[Development-2] #12. Domain-Specific STT Training: Whisper Tiny Fine-Tuning on CPU](https://uno-kim.tistory.com/480)
391
- 15. [[Comparison-8] #13. Vanilla Model vs Custom Fine-Tuned Model (Economics/News Domain)](https://uno-kim.tistory.com/481)
392
-
393
- ---
394
-
395
- ## 🌌 The AMEVA Mobile AI & Automation Ecosystem
396
-
397
- * **🎨 [Termux-Diffusion](https://github.com/uno-km/termux-diffusion)** ([PyPI](https://pypi.org/project/termux-diffusion/) | [npm](https://www.npmjs.com/package/termux-diffusion) | [Docs](https://uno-km.github.io/termux-diffusion/)): Production on-device Stable Diffusion image generation for Android Termux.
398
- * **🌐 [Termux-Playwright](https://github.com/uno-km/termux-playwright-demo)** ([PyPI](https://pypi.org/project/termux-playwright/) | [npm](https://www.npmjs.com/package/termux-playwright) | [Docs](https://uno-km.github.io/termux-playwright-demo/)): Production headless Chromium browser automation & scraping for Android Termux.
399
- * **🧠 [termux-train](https://github.com/uno-km/termux-train)**: On-device Autograd deep learning training & LoRA fine-tuning for Android Termux.
400
- * **🖥️ [AMEVA Workstation Web](https://github.com/uno-km/AMEVA-Workstation-Web)** ([Live Demo](https://ameva-workstation-web-core.vercel.app/)): 100% on-device WebGPU AI workspace & multimedia document intelligence.
401
- * **⚡ [AMEVA-Forge](https://github.com/uno-km/ameva-forge)** ([Docs](https://uno-km.github.io/ameva-forge/)): Real-time WebGPU 3D neural studio & visualization engine.
402
-
403
- ---
404
-
405
- ## ⚖️ Disclaimer (면책 조항)
406
-
407
- > **Disclaimer:**
408
- > *termux-stt is an independent open-source project developed for the Android Termux environment and is not officially affiliated with, endorsed by, or sponsored by the Termux project, OpenAI, or any other third party.*
409
- >
410
- > *(본 프로젝트는 안드로이드 Termux 환경을 위해 개발된 독립적인 오픈소스 라이브러리이며, Termux 공식 프로젝트, OpenAI 및 기타 제3자와 직접적인 제휴 관계가 아닙니다.)*
411
-
412
- ---
413
-
414
- ## 📄 License
415
-
416
- Released under the **MIT License**. Maintained by **uno-km (쌩초보코딩단) / Eunho Kim**.
1
+ # Termux-STT
2
+
3
+ <div align="center">
4
+
5
+ ```
6
+ ████████╗███████╗██████╗ ███╗ ███╗██╗ ██╗██╗ ██╗ ███████╗████████╗████████╗
7
+ ╚══██╔══╝██╔════╝██╔══██╗████╗ ████║██║ ██║╚██╗██╔╝ ██╔════╝╚══██╔══╝╚══██╔══╝
8
+ ██║ █████╗ ██████╔╝██╔████╔██║██║ ██║ ╚███╔╝ █████╗███████╗ ██║ ██║
9
+ ██║ ██╔══╝ ██╔══██╗██║╚██╔╝██║██║ ██║ ██╔██╗ ╚════╝╚════██║ ██║ ██║
10
+ ██║ ███████╗██║ ██║██║ ╚═╝ ██║╚██████╔╝██╔╝ ██╗ ███████║ ██║ ██║
11
+ ╚═╝ ╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝ ╚═╝ ╚═╝
12
+ ```
13
+
14
+ **Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux**
15
+ *Dual-Engine Architecture (Python & Node.js / TypeScript) with Native Bionic ARM64 Acceleration & 0 PyTorch Dependency*
16
+
17
+ <p align="center">
18
+ <a href="https://pypi.org/project/termux-stt/"><img src="https://img.shields.io/pypi/v/termux-stt.svg?style=for-the-badge&color=0088ff&logo=pypi&logoColor=white" alt="PyPI Version" /></a>
19
+ <a href="https://pypi.org/project/termux-stt/"><img src="https://img.shields.io/badge/PyPI%20Downloads-active-0088ff?style=for-the-badge&logo=pypi&logoColor=white" alt="PyPI Downloads" /></a>
20
+ <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/npm/v/termux-stt.svg?style=for-the-badge&color=cb3837&logo=npm&logoColor=white" alt="npm Version" /></a>
21
+ <a href="https://www.npmjs.com/package/termux-stt"><img src="https://img.shields.io/badge/npm%20Downloads-active-cb3837?style=for-the-badge&logo=npm&logoColor=white" alt="npm Downloads" /></a>
22
+ </p>
23
+
24
+ <p align="center">
25
+ <a href="https://uno-km.vercel.app/lib/stt/"><img src="https://img.shields.io/badge/Official_Docs-uno--km.vercel.app%2Flib%2Fstt-004499?style=for-the-badge&logo=vercel&logoColor=white" alt="Live Docs" /></a>
26
+ <a href="https://uno-km.vercel.app/lib/stt/demo.html"><img src="https://img.shields.io/badge/Live_Showcase-▶_Audio_Player-00f5d4?style=for-the-badge&logo=googlechrome&logoColor=0b132b" alt="Live Audio Showcase" /></a>
27
+ <a href="https://github.com/uno-km/termux-stt"><img src="https://img.shields.io/github/stars/uno-km/termux-stt?style=for-the-badge&color=gold&logo=github" alt="GitHub Stars" /></a>
28
+ <a href="https://github.com/uno-km/termux-stt/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg?style=for-the-badge" alt="License" /></a>
29
+ </p>
30
+
31
+ <p align="center">
32
+ <img src="https://img.shields.io/badge/Platform-Android%20Termux%20(ARM64%2Faarch64)-00887A?style=flat-square&logo=android&logoColor=white" alt="Platform" />
33
+ <img src="https://img.shields.io/badge/Engines-whisper.cpp%20%7C%20Vosk%20%7C%20Sherpa--ONNX-38bdf8?style=flat-square" alt="Engines" />
34
+ <img src="https://img.shields.io/badge/Diarization-128d%20X--Vector%20+%20Pure%20Python-a855f7?style=flat-square" alt="Diarization" />
35
+ <img src="https://img.shields.io/badge/RAM-Under%20350MB%20(Tiny%2FBase)-10b981?style=flat-square&logo=shield&logoColor=white" alt="RAM" />
36
+ <img src="https://img.shields.io/badge/Foundation-AOSF_Tier_1-orange?style=flat-square" alt="Foundation" />
37
+ </p>
38
+
39
+ <br/>
40
+
41
+ **[Live Audio Showcase & Demo](https://uno-km.vercel.app/lib/stt/demo.html)** • **[Official Documentation Site (13 Languages)](https://uno-km.vercel.app/lib/stt/)** • **[AMEVA Foundation](https://uno-km.vercel.app/docs/foundation/)** • **[Quickstart](#1-quick-scenario-playbook)** • **[Architecture](#2-why-termux-stt-architectural-pillars)** • **[Benchmarks](#3-empirical-benchmarks-galaxy-a35--exynos-1380)**
42
+
43
+ </div>
44
+
45
+ ---
46
+
47
+ ## AMEVA Foundation — Mobile AI Ecosystem
48
+
49
+ > **"$0 Cloud Egress, 0% External Data Leaks. Transforming every smartphone into an independent on-device AI workstation."**
50
+ >
51
+ > The **AMEVA Open-Source Foundation (AOSF)** builds next-generation, client-centric AI runtimes spanning on-device large models, browser automation, neural network training, and speaker diarization.
52
+
53
+ <div align="center">
54
+
55
+ | Project | Platform & Packages | Core Capability & Technology | Documentation & Demo |
56
+ | :--- | :--- | :--- | :---: |
57
+ | **[termux-stt](https://github.com/uno-km/termux-stt)** | [![PyPI](https://img.shields.io/pypi/v/termux-stt?color=blue&style=flat-square)](https://pypi.org/project/termux-stt/) [![npm](https://img.shields.io/npm/v/termux-stt?color=red&style=flat-square)](https://www.npmjs.com/package/termux-stt) | **Integrated On-Device STT & Pure Python 128d X-Vector Diarization** (Whisper + Vosk + Sherpa) | **[Showcase](https://uno-km.github.io/termux-stt/showcase.html)** |
58
+ | **[termux-diffusion](https://github.com/uno-km/termux-diffusion)** | [![PyPI](https://img.shields.io/pypi/v/termux-diffusion?color=blue&style=flat-square)](https://pypi.org/project/termux-diffusion/) [![npm](https://img.shields.io/npm/v/termux-diffusion?color=red&style=flat-square)](https://www.npmjs.com/package/termux-diffusion) | **Mobile On-Device Stable Diffusion Image Generation** (bfloat16 ARM NEON acceleration) | **[Docs](https://uno-km.github.io/termux-diffusion/)** |
59
+ | **[termux-playwright](https://github.com/uno-km/termux-playwright)** | [![PyPI](https://img.shields.io/pypi/v/termux-playwright?color=blue&style=flat-square)](https://pypi.org/project/termux-playwright/) [![npm](https://img.shields.io/npm/v/termux-playwright?color=red&style=flat-square)](https://www.npmjs.com/package/termux-playwright) | **Non-Root Native Headless Chromium Browser Automation & Scraping** | **[Docs](https://uno-km.github.io/termux-playwright/)** |
60
+ | **[termux-train](https://github.com/uno-km/termux-train)** | [![PyPI](https://img.shields.io/pypi/v/termux-train.svg?color=blue&style=flat-square)](https://pypi.org/project/termux-train/) | **Mobile Native Autograd Neural Network Training & LoRA Fine-Tuning** | **[Docs](https://uno-km.vercel.app/lib/train/)** |
61
+
62
+ </div>
63
+
64
+ ---
65
+
66
+ ## 1. Quick Scenario Playbook
67
+
68
+ ### Installation
69
+
70
+ #### Python SDK:
71
+ ```bash
72
+ pkg update -y && pkg install python ffmpeg git -y
73
+ pip install termux-stt && termux-stt install
74
+ ```
75
+
76
+ #### Node.js / TypeScript:
77
+ ```bash
78
+ pkg update -y && pkg install nodejs-lts ffmpeg git -y
79
+ npm install -g termux-stt && npx termux-stt install
80
+ ```
81
+
82
+ ---
83
+
84
+ ## License
85
+
86
+ Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).
package/package.json CHANGED
@@ -1,50 +1,50 @@
1
- {
2
- "name": "termux-stt",
3
- "version": "1.0.0",
4
- "description": "Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux & Samsung Galaxy",
5
- "main": "index.js",
6
- "types": "index.d.ts",
7
- "bin": {
8
- "termux-stt": "./bin/termux-stt.js"
9
- },
10
- "files": [
11
- "index.js",
12
- "index.d.ts",
13
- "lib",
14
- "bin",
15
- "README.md",
16
- "LICENSE"
17
- ],
18
- "repository": {
19
- "type": "git",
20
- "url": "git+https://github.com/uno-km/termux-stt.git"
21
- },
22
- "keywords": [
23
- "termux",
24
- "stt",
25
- "speech-to-text",
26
- "whisper",
27
- "whisper-cpp",
28
- "vosk",
29
- "sherpa-onnx",
30
- "diarization",
31
- "speaker-diarization",
32
- "android",
33
- "samsung-galaxy",
34
- "edge-ai",
35
- "on-device-ai",
36
- "pure-python",
37
- "exynos",
38
- "snapdragon",
39
- "arm64"
40
- ],
41
- "author": "uno-km (쌩초보코딩단) <hosequelbo@gmail.com>",
42
- "license": "MIT",
43
- "bugs": {
44
- "url": "https://github.com/uno-km/termux-stt/issues"
45
- },
46
- "homepage": "https://uno-km.github.io/termux-stt/",
47
- "engines": {
48
- "node": ">=16.0.0"
49
- }
50
- }
1
+ {
2
+ "name": "termux-stt",
3
+ "version": "1.0.1",
4
+ "description": "Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux & Samsung Galaxy",
5
+ "main": "index.js",
6
+ "types": "index.d.ts",
7
+ "bin": {
8
+ "termux-stt": "./bin/termux-stt.js"
9
+ },
10
+ "files": [
11
+ "index.js",
12
+ "index.d.ts",
13
+ "lib",
14
+ "bin",
15
+ "README.md",
16
+ "LICENSE"
17
+ ],
18
+ "repository": {
19
+ "type": "git",
20
+ "url": "git+https://github.com/uno-km/termux-stt.git"
21
+ },
22
+ "keywords": [
23
+ "termux",
24
+ "stt",
25
+ "speech-to-text",
26
+ "whisper",
27
+ "whisper-cpp",
28
+ "vosk",
29
+ "sherpa-onnx",
30
+ "diarization",
31
+ "speaker-diarization",
32
+ "android",
33
+ "samsung-galaxy",
34
+ "edge-ai",
35
+ "on-device-ai",
36
+ "pure-python",
37
+ "exynos",
38
+ "snapdragon",
39
+ "arm64"
40
+ ],
41
+ "author": "uno-km (AMEVA Foundation) <hosequelbo@gmail.com>",
42
+ "license": "MIT",
43
+ "bugs": {
44
+ "url": "https://github.com/uno-km/termux-stt/issues"
45
+ },
46
+ "homepage": "https://uno-km.github.io/termux-stt/",
47
+ "engines": {
48
+ "node": ">=16.0.0"
49
+ }
50
+ }