termux-stt 1.0.1 β†’ 1.0.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -47,17 +47,18 @@
47
47
  ## AMEVA Foundation β€” Mobile AI Ecosystem
48
48
 
49
49
  > **"$0 Cloud Egress, 0% External Data Leaks. Transforming every smartphone into an independent on-device AI workstation."**
50
- >
51
50
  > The **AMEVA Open-Source Foundation (AOSF)** builds next-generation, client-centric AI runtimes spanning on-device large models, browser automation, neural network training, and speaker diarization.
52
51
 
53
52
  <div align="center">
54
53
 
55
54
  | Project | Platform & Packages | Core Capability & Technology | Documentation & Demo |
56
55
  | :--- | :--- | :--- | :---: |
57
- | **[termux-stt](https://github.com/uno-km/termux-stt)** | [![PyPI](https://img.shields.io/pypi/v/termux-stt?color=blue&style=flat-square)](https://pypi.org/project/termux-stt/) [![npm](https://img.shields.io/npm/v/termux-stt?color=red&style=flat-square)](https://www.npmjs.com/package/termux-stt) | **Integrated On-Device STT & Pure Python 128d X-Vector Diarization** (Whisper + Vosk + Sherpa) | **[Showcase](https://uno-km.github.io/termux-stt/showcase.html)** |
58
- | **[termux-diffusion](https://github.com/uno-km/termux-diffusion)** | [![PyPI](https://img.shields.io/pypi/v/termux-diffusion?color=blue&style=flat-square)](https://pypi.org/project/termux-diffusion/) [![npm](https://img.shields.io/npm/v/termux-diffusion?color=red&style=flat-square)](https://www.npmjs.com/package/termux-diffusion) | **Mobile On-Device Stable Diffusion Image Generation** (bfloat16 ARM NEON acceleration) | **[Docs](https://uno-km.github.io/termux-diffusion/)** |
59
- | **[termux-playwright](https://github.com/uno-km/termux-playwright)** | [![PyPI](https://img.shields.io/pypi/v/termux-playwright?color=blue&style=flat-square)](https://pypi.org/project/termux-playwright/) [![npm](https://img.shields.io/npm/v/termux-playwright?color=red&style=flat-square)](https://www.npmjs.com/package/termux-playwright) | **Non-Root Native Headless Chromium Browser Automation & Scraping** | **[Docs](https://uno-km.github.io/termux-playwright/)** |
60
- | **[termux-train](https://github.com/uno-km/termux-train)** | [![PyPI](https://img.shields.io/pypi/v/termux-train.svg?color=blue&style=flat-square)](https://pypi.org/project/termux-train/) | **Mobile Native Autograd Neural Network Training & LoRA Fine-Tuning** | **[Docs](https://uno-km.vercel.app/lib/train/)** |
56
+ | πŸŽ™οΈ **[termux-stt](https://github.com/uno-km/termux-stt)** | [![Open Collective](https://img.shields.io/badge/Open_Collective-AOSF_Fund-004499?style=flat&logo=opencollective)](https://opencollective.com/ameva-fund) [![GitHub Sponsors](https://img.shields.io/badge/GitHub_Sponsors-uno--km-ea4aaa?style=flat&logo=githubsponsors)](https://github.com/sponsors/uno-km)<br/>[![PyPI](https://img.shields.io/pypi/v/termux-stt?color=blue&style=flat-square)](https://pypi.org/project/termux-stt/) [![npm](https://img.shields.io/npm/v/termux-stt?color=red&style=flat-square)](https://www.npmjs.com/package/termux-stt) | **Integrated On-Device STT & Pure Python 128d X-Vector Diarization** (Whisper + Vosk + Sherpa) | **[Showcase](https://uno-km.github.io/termux-stt/showcase.html)** β€’ **[Docs](https://uno-km.vercel.app/lib/stt/)** |
57
+ | 🎨 **[termux-diffusion](https://github.com/uno-km/termux-diffusion)** | [![PyPI](https://img.shields.io/pypi/v/termux-diffusion?color=blue&style=flat-square)](https://pypi.org/project/termux-diffusion/) [![npm](https://img.shields.io/npm/v/termux-diffusion?color=red&style=flat-square)](https://www.npmjs.com/package/termux-diffusion) | **Mobile On-Device Stable Diffusion Image Generation** (bfloat16 ARM NEON acceleration) | **[Docs](https://uno-km.github.io/termux-diffusion/)** |
58
+ | 🌐 **[termux-playwright](https://github.com/uno-km/termux-playwright)** | [![PyPI](https://img.shields.io/pypi/v/termux-playwright?color=blue&style=flat-square)](https://pypi.org/project/termux-playwright/) [![npm](https://img.shields.io/npm/v/termux-playwright?color=red&style=flat-square)](https://www.npmjs.com/package/termux-playwright) | **Non-Root Native Headless Chromium Browser Automation & Scraping** | **[Docs](https://uno-km.github.io/termux-playwright/)** |
59
+ | 🧠 **[termux-train](https://github.com/uno-km/termux-train)** | [![PyPI](https://img.shields.io/pypi/v/termux-train.svg?color=blue&style=flat-square)](https://pypi.org/project/termux-train/) | **Mobile Native Autograd Neural Network Training & LoRA Fine-Tuning** | **[Docs](https://uno-km.vercel.app/lib/train/)** |
60
+ | πŸ–₯️ **[AMEVA Workstation](https://github.com/uno-km/AMEVA-Workstation-Web)** | [![WebGPU](https://img.shields.io/badge/WebGPU-100%25_On--Device-00f5d4?style=flat-square)](https://ameva-workstation-web-core.vercel.app/) | **100% Client-Side WebGPU Multimodal Document Intelligence Workspace** | **[Live Demo](https://ameva-workstation-web-core.vercel.app/)** |
61
+ | ⚑ **[AMEVA-Forge](https://github.com/uno-km/ameva-forge)** | [![WebGPU](https://img.shields.io/badge/3D_Studio-WebGPU-purple?style=flat-square)](https://uno-km.github.io/ameva-forge/demo.html) | **Real-Time 3D Neural Studio & WebGPU Visualization Engine** | **[Live Demo](https://uno-km.github.io/ameva-forge/demo.html)** |
61
62
 
62
63
  </div>
63
64
 
@@ -76,11 +77,289 @@ pip install termux-stt && termux-stt install
76
77
  #### Node.js / TypeScript:
77
78
  ```bash
78
79
  pkg update -y && pkg install nodejs-lts ffmpeg git -y
79
- npm install -g termux-stt && npx termux-stt install
80
+ ---
81
+
82
+ ## 1. Quick Scenario Playbook
83
+
84
+ ### 1-Click Installation
85
+
86
+ #### Python SDK:
87
+ ```bash
88
+ pip install termux-stt && termux-stt install
89
+ ```
90
+
91
+ #### Node.js / TypeScript:
92
+ ```bash
93
+ npm install -g termux-stt && termux-stt install
94
+ ```
95
+
96
+ ---
97
+
98
+ ### πŸŽ™οΈ Try with the Included Sample Audio! (JFK 1-Minute Speech)
99
+
100
+ `termux-stt` includes John F. Kennedy's 1961 Inaugural Address (60.00s 16kHz Mono PCM) in `samples/jfk_1min.wav` for instant out-of-the-box testing.
101
+
102
+ #### Option A: CLI One-Liner (Python / npm)
103
+ ```bash
104
+ # 1. Transcribe the 60s sample speech with Whisper (auto-generates SRT subtitles)
105
+ termux-stt transcribe samples/jfk_1min.wav --engine whisper --model tiny --format srt
106
+
107
+ # 2. Transcribe and separate speakers (128d X-Vector Diarization)
108
+ termux-stt diarize samples/jfk_1min.wav --speakers 2 --format text
109
+
110
+ # 3. Benchmark on-device latency & RTF
111
+ termux-stt benchmark --audio samples/jfk_1min.wav --model tiny
112
+ ```
113
+
114
+ #### Option B: Python SDK
115
+ ```python
116
+ from termux_stt import create_engine
117
+
118
+ # 1. Initialize Whisper Engine (auto-loads native ARM NEON binary)
119
+ engine = create_engine("whisper", model="tiny", lang="en", threads=4)
120
+
121
+ # 2. Transcribe JFK 60s speech
122
+ result = engine.transcribe("samples/jfk_1min.wav")
123
+
124
+ print("Transcript:\n", result.text)
125
+ print("SRT Subtitles:\n", result.to_srt())
126
+ ```
127
+
128
+ #### Option C: Node.js / TypeScript
129
+ ```javascript
130
+ const { createEngine } = require("termux-stt");
131
+
132
+ async function main() {
133
+ const engine = createEngine("whisper", { model: "tiny", lang: "en", threads: 4 });
134
+ const result = await engine.transcribe("samples/jfk_1min.wav");
135
+
136
+ console.log("Transcript:", result.text);
137
+ console.log("SRT Subtitles:\n", result.toSrt());
138
+ }
139
+ main();
140
+ ```
141
+
142
+ ---
143
+
144
+ ## 2. Advanced Scenarios
145
+
146
+ ### [Streaming] Scenario 3: Real-Time Microphone Streaming
147
+
148
+ Speak into your smartphone microphone and receive real-time transcribed text with sub-second latency:
149
+
150
+ ```python
151
+ from termux_stt import create_engine
152
+
153
+ # Use ultra-lightweight Tiny model (RTF 0.80 on Exynos 1380)
154
+ engine = create_engine("whisper", model="tiny", lang="ko")
155
+
156
+ print("πŸŽ™οΈ Listening... Speak into your phone microphone (Ctrl+C to stop)")
157
+ for segment in engine.stream_mic():
158
+ print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")
80
159
  ```
81
160
 
82
161
  ---
83
162
 
84
- ## License
163
+ ### [Diarize] Scenario 4: Hybrid Speaker Diarization ("Who Spoke When?")
164
+
165
+ Run high-precision speaker diarization without PyTorch or CUDA:
166
+
167
+ ```python
168
+ from termux_stt import create_engine
169
+
170
+ # Hybrid Pipeline: Vosk 128d X-Vector + Whisper STT + Pure Python K-Means
171
+ engine = create_engine("hybrid", lang="ko", num_speakers=2)
172
+ result = engine.diarize("interview.wav")
173
+
174
+ for seg in result.segments:
175
+ print(f"[{seg.speaker}] ({seg.start:.1f}s - {seg.end:.1f}s): {seg.text}")
176
+ ```
177
+
178
+ *Output Example:*
179
+ ```text
180
+ [Speaker_0] (0.0s - 3.5s): 였늘 경제 λΈŒλ¦¬ν•‘μ„ μ‹œμž‘ν•˜κ² μŠ΅λ‹ˆλ‹€.
181
+ [Speaker_1] (3.8s - 7.2s): λ„€, 였늘 μ½”μŠ€ν”Ό μ§€μˆ˜κ°€ 외ꡭ인 순맀수둜 μƒμŠΉ λ§ˆκ°ν–ˆμŠ΅λ‹ˆλ‹€.
182
+ [Speaker_0] (7.5s - 10.1s): λ°˜λ„μ²΄ μ„Ήν„° 동ν–₯은 μ–΄λ–€κ°€μš”?
183
+ ```
184
+
185
+ ---
186
+
187
+ ## 2. πŸ›οΈ Why termux-stt? Architectural Pillars
188
+
189
+ ```
190
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
191
+ β”‚ termux-stt Architecture β”‚
192
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
193
+ β”‚ User Interface Layer β”‚ Engine Abstraction β”‚ Output / Export Layer β”‚
194
+ β”‚ β€’ Python API β”‚ β€’ EngineRegistry β”‚ β€’ JSON / SRT / VTT β”‚
195
+ β”‚ β€’ Node.js API β”‚ β€’ create_engine() β”‚ β€’ RTTM (Diarization) β”‚
196
+ β”‚ β€’ CLI (termux-stt) β”‚ β€’ ModelHub & Cache β”‚ β€’ Streaming Callback β”‚
197
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
198
+ β”‚ Core Pipeline Layer β”‚
199
+ β”‚ Audio Loader (7 Formats) βž” Preprocessor (16kHz Mono) βž” Silero-VAD Filter β”‚
200
+ β”‚ βž” Multi-Engine STT (whisper.cpp / Vosk / Sherpa-ONNX) β”‚
201
+ β”‚ βž” Hybrid Diarizer (128d X-Vector βž” Pure Python Cosine/K-Means βž” Time Align)β”‚
202
+ β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
203
+ β”‚ Platform & Process Isolation β”‚
204
+ β”‚ β€’ Subprocess Isolation (Host Python never crashes on C++ Segfault) β”‚
205
+ β”‚ β€’ MobileGuard (WakeLock, Doze Mode Bypass, Phantom Process Killer Shield) β”‚
206
+ β”‚ β€’ Bionic ARM64 NEON & FP16 SIMD Acceleration β”‚
207
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
208
+ ```
209
+
210
+ 1. **Subprocess Isolation**: C++ inference runs in isolated native subprocesses. Memory errors or Segfaults never kill the host Python/Node.js application.
211
+ 2. **Pure Python Clustering**: Cosine distance matrix and K-Means clustering are implemented with 0 external dependencies (`numpy` / `scikit-learn` optional, not required).
212
+ 3. **Automated Android Bionic Fixes**: Solves `sys.platform = 'linux'` spoofing, libvosk CFFI extraction, and ffmpeg audio format normalization (`16kHz / 1ch / PCM s16le`) under the hood.
213
+ 4. **Mobile Battery & CPU Guard**: Manages Android WakeLocks and background task states to prevent process termination when the screen locks.
214
+
215
+ ---
216
+
217
+ ## 3. πŸ“Š Empirical Benchmarks (Galaxy A35 / Exynos 1380)
218
+
219
+ > *Measured on Samsung Galaxy A35 5G (Exynos 1380 4x A78 + 4x A55, 6GB RAM, Android 14 Termux).*
220
+
221
+ | Engine / Pipeline | Model | Peak RAM | RTF (Speed) | KO Accuracy | Diarization Support | Termux Rating |
222
+ | :--- | :--- | :---: | :---: | :---: | :---: | :---: |
223
+ | **whisper.cpp** | `ggml-tiny` (39M) | **~150 MB** | **0.80** | 85% | ❌ External | ⭐⭐⭐⭐⭐ |
224
+ | **whisper.cpp** | `ggml-base` (74M) | ~250 MB | 1.20 | 88% | ❌ External | ⭐⭐⭐⭐ |
225
+ | **whisper.cpp** | `ggml-medium` (769M) | ~1.5 GB | 3.40 | **95%+** | ❌ External | ⭐⭐⭐ (Golden Acc) |
226
+ | **Vosk** | `small-ko-0.22` (42M) | **~100 MB** | **0.25** | 78% | βœ… 128d X-Vector | ⭐⭐⭐ |
227
+ | **Sherpa-ONNX** | `Zipformer` | ~300 MB | **0.42** | 86% | βœ… CAM++ | ⭐⭐⭐⭐ |
228
+ | **Pyannote.audio 3.1** | `diarization-3.1` | **> 3.5 GB** | 2.80~3.50 | N/A | βœ… Gold Standard | ❌ OOM Crashes |
229
+ | **termux-stt (Hybrid)** | `Vosk + Whisper Base`| **~350 MB** | **1.45** | **92%+** | **βœ… Built-in K-Means** | **⭐⭐⭐⭐⭐ (Recommended)** |
230
+
231
+ ---
232
+
233
+ ## 4. βš™οΈ Engine Comparison Matrix
234
+
235
+ | Feature | `whisper.cpp` | `Vosk` | `Sherpa-ONNX` | `Hybrid (Vosk+Whisper)` |
236
+ | :--- | :---: | :---: | :---: | :---: |
237
+ | **Primary Strength** | Highest Text Accuracy | Ultra-Low RAM & Fast | Ultra-Low Latency | **STT + Diarization Combined** |
238
+ | **Memory Footprint** | 150MB ~ 1.5GB | **< 100MB** | 300MB ~ 500MB | **~ 350MB** |
239
+ | **Real-Time Factor (RTF)** | 0.80 (Tiny) | **0.25 (Blazing)** | **0.42 (Fast)** | 1.45 (Full Pipeline) |
240
+ | **Speaker Diarization** | ❌ None | ⚠️ Basic X-Vector | ⚠️ CAM++ C++ | **βœ… High-Precision Aligned** |
241
+ | **Recommended Use Case** | Quality Transcripts | Embedded / Low Spec | Live Voice Assistant | **Meetings / Interviews** |
242
+
243
+ ---
244
+
245
+ ## 5. πŸ“š Complete API Reference Summary
246
+
247
+ ### Python API
248
+
249
+ ```python
250
+ import termux_stt
251
+
252
+ # Create Engine with fine-grained control parameters
253
+ engine = termux_stt.create_engine(
254
+ engine="whisper", # "whisper" | "vosk" | "sherpa" | "hybrid"
255
+ model="base", # "tiny" | "base" | "small" | "medium" | "custom"
256
+ lang="ko", # ISO 639-1 language code
257
+ num_speakers=0, # 0 = disabled, 2+ = enable diarization
258
+ threads=4, # CPU threads (defaults to big cores count)
259
+ vad=True, # Enable Silero-VAD silence stripping
260
+ quantization="q5_1", # "f16" | "q8_0" | "q5_1" | "q4_0"
261
+ prompt="경제 λΈŒλ¦¬ν•‘", # Initial decoding context / vocabulary
262
+ beam_size=5, # Beam search beam size
263
+ temperature=0.0 # Sampling temperature
264
+ )
265
+
266
+ # Transcribe File
267
+ result = engine.transcribe("audio.wav")
268
+ # Returns: TranscriptResult(text=str, segments=List[Segment], language=str, duration=float)
269
+
270
+ # Export Methods
271
+ result.to_json() # Structured JSON string
272
+ result.to_srt() # Standard SRT subtitle format
273
+ result.to_vtt() # WebVTT subtitle format
274
+ result.to_rttm() # NIST RTTM diarization format
275
+
276
+ # Stream Microphone
277
+ for seg in engine.stream_mic(duration=30.0):
278
+ print(f"[{seg.speaker}] {seg.text}")
279
+
280
+ # Speaker Diarization
281
+ diar_result = engine.diarize("meeting.wav", num_speakers=2)
282
+ ```
283
+
284
+ ### CLI Reference
285
+
286
+ ```bash
287
+ # General Syntax
288
+ termux-stt [COMMAND] [OPTIONS] [FILE]
289
+
290
+ # Commands
291
+ termux-stt transcribe [FILE] # Transcribe audio file (--prompt, --beam-size, --translate)
292
+ termux-stt listen # Real-time microphone transcription
293
+ termux-stt diarize [FILE] # Perform speaker diarization
294
+ termux-stt models list # List installed and available models
295
+ termux-stt models download [M] # Download specific model
296
+ termux-stt doctor # Run hardware and environment diagnostics
297
+ termux-stt benchmark # Run performance benchmark suite
298
+ ```
299
+
300
+ ---
301
+
302
+ ## 6. πŸ› οΈ Troubleshooting & Android FAQs
303
+
304
+ ### Q1: `pip install vosk` fails with CMake or wheel error on Android
305
+ * **Cause**: Vosk does not publish official prebuilt aarch64-android wheels on PyPI.
306
+ * **Solution**: `termux-stt-install` automatically extracts `libvosk.so` from the official Android AAR and generates the CFFI bindings.
307
+
308
+ ### Q2: Whisper crashes on 44.1kHz stereo MP3/M4A files
309
+ * **Cause**: Whisper models strictly require single-channel 16,000Hz 16-bit PCM WAV.
310
+ * **Solution**: `termux-stt` automatically runs `ffmpeg` normalization on any audio format (`mp3`, `m4a`, `flac`, `ogg`, `opus`, `webm`).
311
+
312
+ ### Q3: Process killed after 10 minutes in background
313
+ * **Cause**: Android Phantom Process Killer terminates background tasks.
314
+ * **Solution**: Enable Termux WakeLock (`termux-wake-lock`) and disable battery optimization for Termux in Android Settings.
315
+
316
+ ---
317
+
318
+ ## 7. πŸ” 15-Part Empirical Research Blog Series
319
+
320
+ This framework is built upon the exhaustive 15-part research series published on [Eunho Kim's Technical Blog (μš°λ…Έν‚΄ ν‹°μŠ€ν† λ¦¬)](https://uno-kim.tistory.com/):
321
+
322
+ 1. [[Whisper.cpp] #1. Edge Agent AI: Whisper.cpp Speech Processing (Base vs Tiny)](https://uno-kim.tistory.com/467)
323
+ 2. [[Audio Extraction] #2. Extracting Specific Audio Segments on Android](https://uno-kim.tistory.com/468)
324
+ 3. [[Whisper.cpp] #3. Korean STT Conversion Comparison (4 Models)](https://uno-kim.tistory.com/469)
325
+ 4. [[Comparison-1] #4. STT + Speaker Diarization: 3 Lightweight Engines + Pyannote](https://uno-kim.tistory.com/472)
326
+ 5. [[Comparison-2] #5. Sherpa-ONNX Execution, Diarization, and Troubleshooting](https://uno-kim.tistory.com/473)
327
+ 6. [[Comparison-3] #6. Speaker Diarization using Pyannote Model](https://uno-kim.tistory.com/471)
328
+ 7. [[Pyannote] Troubleshooting Diarization in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/470)
329
+ 8. [[Comparison-4] #7. Vosk Execution, Speaker Diarization, and Troubleshooting](https://uno-kim.tistory.com/475)
330
+ 9. [[Vosk] Troubleshooting Vosk in Mobile/Termux/ARM Environments](https://uno-kim.tistory.com/474)
331
+ 10. [[Comparison-5] #8. Vosk / Pyannote / Sherpa-ONNX / Whisper.cpp Comprehensive Comparison](https://uno-kim.tistory.com/476)
332
+ 11. [[Comparison-6] #9. Vosk + Whisper.cpp Hybrid Pipeline & X-Vector Diarization](https://uno-kim.tistory.com/477)
333
+ 12. [[Development-1] #10. Large-Scale Batch Automation & Task Management Architecture](https://uno-kim.tistory.com/478)
334
+ 13. [[Comparison-7] #11. STT + Diarization Final: Small vs Turbo & Optimization Magic (4-Model Benchmark)](https://uno-kim.tistory.com/479)
335
+ 14. [[Development-2] #12. Domain-Specific STT Training: Whisper Tiny Fine-Tuning on CPU](https://uno-kim.tistory.com/480)
336
+ 15. [[Comparison-8] #13. Vanilla Model vs Custom Fine-Tuned Model (Economics/News Domain)](https://uno-kim.tistory.com/481)
337
+
338
+ ---
339
+
340
+ ## βš–οΈ Disclaimer (λ©΄μ±… μ‘°ν•­)
341
+
342
+ > **Disclaimer:**
343
+ > *termux-stt is an independent open-source project developed for the Android Termux environment and is not officially affiliated with, endorsed by, or sponsored by the Termux project, OpenAI, or any other third party.*
344
+ >
345
+ > *(λ³Έ ν”„λ‘œμ νŠΈλŠ” μ•ˆλ“œλ‘œμ΄λ“œ Termux ν™˜κ²½μ„ μœ„ν•΄ 개발된 독립적인 μ˜€ν”ˆμ†ŒμŠ€ 라이브러리이며, Termux 곡식 ν”„λ‘œμ νŠΈ, OpenAI 및 기타 제3μžμ™€ 직접적인 제휴 관계가 μ•„λ‹™λ‹ˆλ‹€.)*
346
+
347
+ ---
348
+
349
+ ## πŸ“„ License
350
+
351
+ Released under the **MIT License**. Maintained by **AMEVA Foundation & uno-km (μŒ©μ΄ˆλ³΄μ½”λ”©λ‹¨) / Eunho Kim**.
352
+
353
+
354
+ ---
355
+
356
+ ## πŸ’– Sponsorship & Community Backing
357
+
358
+ AMEVA is an independent open-source public good governed under the **AMEVA Open-Source Foundation (AOSF)**. All sponsorship funds are 100% publicly audited and dedicated to physical ARM64 testbeds and CI/CD GPU runners.
85
359
 
360
+ - **Open Collective (Non-Profit 501(c)(6))**: [https://opencollective.com/ameva-fund](https://opencollective.com/ameva-fund)
361
+ - **GitHub Sponsors**: [https://github.com/sponsors/uno-km](https://github.com/sponsors/uno-km)
362
+ - **Official Foundation Portal**: [https://uno-km.vercel.app/docs/foundation/sponsorship.html](https://uno-km.vercel.app/docs/foundation/sponsorship.html)
363
+ =======
86
364
  Apache License 2.0. Copyright (c) 2026 uno-km (AMEVA Foundation).
365
+ >>>>>>> Stashed changes
package/bin/termux-stt.js CHANGED
@@ -4,13 +4,25 @@
4
4
  * termux-stt CLI for Node.js / npx
5
5
  */
6
6
 
7
- const { spawn } = require('child_process');
7
+ const { spawn, spawnSync } = require('child_process');
8
8
 
9
9
  const args = process.argv.slice(2);
10
10
 
11
- // Forward to Python CLI
11
+ function getPythonExecutable() {
12
+ const candidates = ['python3', 'python'];
13
+ for (const cmd of candidates) {
14
+ try {
15
+ const res = spawnSync(cmd, ['--version'], { stdio: 'ignore' });
16
+ if (res.status === 0) return cmd;
17
+ } catch (_) {}
18
+ }
19
+ return 'python3';
20
+ }
21
+
22
+ const pythonExe = getPythonExecutable();
12
23
  const pythonArgs = ['-m', 'termux_stt.cli.main', ...args];
13
- const proc = spawn('python', pythonArgs, {
24
+
25
+ const proc = spawn(pythonExe, pythonArgs, {
14
26
  stdio: 'inherit',
15
27
  env: process.env
16
28
  });
package/package.json CHANGED
@@ -1,17 +1,19 @@
1
1
  {
2
2
  "name": "termux-stt",
3
- "version": "1.0.1",
3
+ "version": "1.0.9",
4
4
  "description": "Production-Grade On-Device Speech-to-Text & Speaker Diarization Framework for Android Termux & Samsung Galaxy",
5
5
  "main": "index.js",
6
6
  "types": "index.d.ts",
7
7
  "bin": {
8
- "termux-stt": "./bin/termux-stt.js"
8
+ "termux-stt": "./bin/termux-stt.js",
9
+ "termux-stt-install": "./bin/termux-stt.js"
9
10
  },
10
11
  "files": [
11
12
  "index.js",
12
13
  "index.d.ts",
13
14
  "lib",
14
15
  "bin",
16
+ "samples",
15
17
  "README.md",
16
18
  "LICENSE"
17
19
  ],
Binary file