termux-stt 1.2.2 → 1.2.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -21,13 +21,16 @@ pkg update -y
21
21
  pkg install -y clang python python-numpy nodejs termux-api ffmpeg pulseaudio
22
22
  ```
23
23
 
24
- ### 1.2 Python SDK & Global CLI Installation
24
+ ### 1.2 Pure Package Installation (Zero-Compilation Bundled Binary)
25
25
  Install the core package from PyPI via `pip`:
26
26
  ```bash
27
27
  pip install --upgrade pip
28
28
  pip install termux-stt
29
29
  ```
30
30
 
31
+ > [!NOTE]
32
+ > **Bundled ARM64 Binary Architecture**: The official Python universal wheel (`termux_stt-*.whl`) directly bundles a pre-compiled Android ARM64 (Bionic libc) `whisper-cli` ELF executable inside `termux_stt/bin/whisper-cli`. When installed via `pip`, this binary is automatically unpacked into Python `site-packages`. Pure 16kHz WAV transcription is functional immediately without requiring any local C/C++ compiler toolchain.
33
+
31
34
  To install with development extras:
32
35
  ```bash
33
36
  pip install "termux-stt[dev]"
@@ -36,19 +39,44 @@ pip install "termux-stt[dev]"
36
39
  ### 1.3 Node.js / TypeScript SDK & CLI Installation
37
40
  Install globally or locally via `npm`:
38
41
  ```bash
39
- # Global CLI installation
42
+ # Global CLI installation (bridges to underlying Python runtime)
40
43
  npm install -g termux-stt
41
44
 
42
45
  # Project dependency installation
43
46
  npm install termux-stt
44
47
  ```
45
48
 
46
- ### 1.4 Zero-Compilation 1-Click Setup
47
- Termux-STT bundles precompiled ARM64 native binaries (`whisper-cli`) and automated model provisioners. Provision the environment with a single command:
49
+ ### 1.4 Post-Installation Automated Environment Provisioner (`termux-stt install`)
50
+ While pure package installation provides instant offline WAV inference, production deployments involving compressed media (MP3/M4A/FLAC), live microphone streaming, or mobile GPU acceleration require full environment provisioning. Run the automated 1-click provisioner:
51
+
48
52
  ```bash
49
53
  termux-stt-install
54
+ # Or equivalently:
55
+ termux-stt install
50
56
  ```
51
57
 
58
+ The automated engine installer (`EngineInstaller`) executes a 3-stage provisioning pipeline:
59
+ 1. **Native System Dependencies (`install_system_dependencies`)**:
60
+ Automatically invokes Termux `pkg` to install `ffmpeg`, `libbluray`, `libxml2`, `git`, `termux-api`, and `curl`, enabling universal audio decoding and microphone capture via Android APIs.
61
+ 2. **Adaptive Engine Binary Provisioning (`install_whisper_cpp`)**:
62
+ - **Vulkan GPU Silicon Detected**: Inspects `/system/lib64/libvulkan.so` and `Doctor().quick_probe()`. If mobile GPU compute is available, provisions `cmake`, `make`, `clang`, clones `whisper.cpp`, and compiles a device-tailored native binary with `-DGGML_VULKAN=ON` and `-DVulkan_LIBRARY=/system/lib64/libvulkan.so`, installing it to `$PREFIX/bin/whisper-cli` and `$HOME/.local/bin/whisper-cli` with top execution priority.
63
+ - **CPU-Only Fallback**: If Vulkan is absent, downloads the pre-built optimized ARM64 NEON static binary directly from GitHub Releases to `$HOME/.local/bin/whisper-cli`.
64
+ 3. **Sub-Engine Ecosystem Provisioning (`install_vosk`, `install_sherpa_onnx`)**:
65
+ Provisions `vosk` for sub-30ms real-time streaming and `sherpa-onnx` for next-generation ONNX Zipformer models, while pre-initializing model cache structures in `~/.cache/termux-stt/models/`.
66
+
67
+ ### 1.5 Pure Install vs. Post-Install Comparison Matrix
68
+
69
+ | Feature / Capability | Pure Install (`pip install termux-stt`) | Post-Install (`termux-stt install`) |
70
+ | :--- | :--- | :--- |
71
+ | **Native Binary State** | Bundled CPU-NEON static binary (`site-packages/termux_stt/bin/`) | Dynamic: Local Vulkan GPU compilation or updated ARM64 release |
72
+ | **GPU Acceleration** | Hardware binding attempted via ameva-runtime | Fully compiled with native SPIR-V Vulkan shaders (`-DGGML_VULKAN=ON`) |
73
+ | **Supported Audio Formats** | Uncompressed WAV (16kHz PCM) | Universal: MP3, M4A, AAC, FLAC, OGG, WAV (via system `ffmpeg`) |
74
+ | **Live Microphone Stream** | Requires manual `termux-api` installation | Automated `termux-api` package provisioning |
75
+ | **Streaming Engine (Vosk)** | Skipped (Whisper-only mode) | Automated `vosk` pip package & model cache configuration |
76
+ | **Zipformer (Sherpa-ONNX)** | Skipped | Automated `sherpa-onnx` pip package & model cache configuration |
77
+ | **Setup Time** | Instantaneous (~3–5 seconds) | ~1–3 minutes (depending on whether local compilation occurs) |
78
+ | **Recommended Use Case** | Quick smoke testing, batch WAV inference | Production services, 24/7 background daemons, mobile GPU offloading |
79
+
52
80
  ---
53
81
 
54
82
  ## 2. GPU Hardware Acceleration Provisioning (`ameva-runtime`)
package/README.pypi.md CHANGED
@@ -21,13 +21,16 @@ pkg update -y
21
21
  pkg install -y clang python python-numpy nodejs termux-api ffmpeg pulseaudio
22
22
  ```
23
23
 
24
- ### 1.2 Python SDK & Global CLI Installation
24
+ ### 1.2 Pure Package Installation (Zero-Compilation Bundled Binary)
25
25
  Install the core package from PyPI via `pip`:
26
26
  ```bash
27
27
  pip install --upgrade pip
28
28
  pip install termux-stt
29
29
  ```
30
30
 
31
+ > [!NOTE]
32
+ > **Bundled ARM64 Binary Architecture**: The official Python universal wheel (`termux_stt-*.whl`) directly bundles a pre-compiled Android ARM64 (Bionic libc) `whisper-cli` ELF executable inside `termux_stt/bin/whisper-cli`. When installed via `pip`, this binary is automatically unpacked into Python `site-packages`. Pure 16kHz WAV transcription is functional immediately without requiring any local C/C++ compiler toolchain.
33
+
31
34
  To install with development extras:
32
35
  ```bash
33
36
  pip install "termux-stt[dev]"
@@ -36,19 +39,44 @@ pip install "termux-stt[dev]"
36
39
  ### 1.3 Node.js / TypeScript SDK & CLI Installation
37
40
  Install globally or locally via `npm`:
38
41
  ```bash
39
- # Global CLI installation
42
+ # Global CLI installation (bridges to underlying Python runtime)
40
43
  npm install -g termux-stt
41
44
 
42
45
  # Project dependency installation
43
46
  npm install termux-stt
44
47
  ```
45
48
 
46
- ### 1.4 Zero-Compilation 1-Click Setup
47
- Termux-STT bundles precompiled ARM64 native binaries (`whisper-cli`) and automated model provisioners. Provision the environment with a single command:
49
+ ### 1.4 Post-Installation Automated Environment Provisioner (`termux-stt install`)
50
+ While pure package installation provides instant offline WAV inference, production deployments involving compressed media (MP3/M4A/FLAC), live microphone streaming, or mobile GPU acceleration require full environment provisioning. Run the automated 1-click provisioner:
51
+
48
52
  ```bash
49
53
  termux-stt-install
54
+ # Or equivalently:
55
+ termux-stt install
50
56
  ```
51
57
 
58
+ The automated engine installer (`EngineInstaller`) executes a 3-stage provisioning pipeline:
59
+ 1. **Native System Dependencies (`install_system_dependencies`)**:
60
+ Automatically invokes Termux `pkg` to install `ffmpeg`, `libbluray`, `libxml2`, `git`, `termux-api`, and `curl`, enabling universal audio decoding and microphone capture via Android APIs.
61
+ 2. **Adaptive Engine Binary Provisioning (`install_whisper_cpp`)**:
62
+ - **Vulkan GPU Silicon Detected**: Inspects `/system/lib64/libvulkan.so` and `Doctor().quick_probe()`. If mobile GPU compute is available, provisions `cmake`, `make`, `clang`, clones `whisper.cpp`, and compiles a device-tailored native binary with `-DGGML_VULKAN=ON` and `-DVulkan_LIBRARY=/system/lib64/libvulkan.so`, installing it to `$PREFIX/bin/whisper-cli` and `$HOME/.local/bin/whisper-cli` with top execution priority.
63
+ - **CPU-Only Fallback**: If Vulkan is absent, downloads the pre-built optimized ARM64 NEON static binary directly from GitHub Releases to `$HOME/.local/bin/whisper-cli`.
64
+ 3. **Sub-Engine Ecosystem Provisioning (`install_vosk`, `install_sherpa_onnx`)**:
65
+ Provisions `vosk` for sub-30ms real-time streaming and `sherpa-onnx` for next-generation ONNX Zipformer models, while pre-initializing model cache structures in `~/.cache/termux-stt/models/`.
66
+
67
+ ### 1.5 Pure Install vs. Post-Install Comparison Matrix
68
+
69
+ | Feature / Capability | Pure Install (`pip install termux-stt`) | Post-Install (`termux-stt install`) |
70
+ | :--- | :--- | :--- |
71
+ | **Native Binary State** | Bundled CPU-NEON static binary (`site-packages/termux_stt/bin/`) | Dynamic: Local Vulkan GPU compilation or updated ARM64 release |
72
+ | **GPU Acceleration** | Hardware binding attempted via ameva-runtime | Fully compiled with native SPIR-V Vulkan shaders (`-DGGML_VULKAN=ON`) |
73
+ | **Supported Audio Formats** | Uncompressed WAV (16kHz PCM) | Universal: MP3, M4A, AAC, FLAC, OGG, WAV (via system `ffmpeg`) |
74
+ | **Live Microphone Stream** | Requires manual `termux-api` installation | Automated `termux-api` package provisioning |
75
+ | **Streaming Engine (Vosk)** | Skipped (Whisper-only mode) | Automated `vosk` pip package & model cache configuration |
76
+ | **Zipformer (Sherpa-ONNX)** | Skipped | Automated `sherpa-onnx` pip package & model cache configuration |
77
+ | **Setup Time** | Instantaneous (~3–5 seconds) | ~1–3 minutes (depending on whether local compilation occurs) |
78
+ | **Recommended Use Case** | Quick smoke testing, batch WAV inference | Production services, 24/7 background daemons, mobile GPU offloading |
79
+
52
80
  ---
53
81
 
54
82
  ## 2. GPU Hardware Acceleration Provisioning (`ameva-runtime`)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "termux-stt",
3
- "version": "1.2.2",
3
+ "version": "1.2.3",
4
4
  "description": "On-device Speech-to-Text & speaker diarization framework utilizing device resources for Android Termux",
5
5
  "main": "index.js",
6
6
  "types": "index.d.ts",