termux-vision 1.4.6 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (5) hide show
  1. package/README.md +304 -319
  2. package/README.pypi.md +132 -121
  3. package/bin/cli.js +51 -51
  4. package/lib/vlm.js +349 -349
  5. package/package.json +55 -55
package/README.md CHANGED
@@ -1,319 +1,304 @@
1
- # Termux-Vision: On-Device Computer Vision & Multimodal VLM Framework
2
-
3
- [![PyPI](https://img.shields.io/pypi/v/termux-vision.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-vision/)
4
- [![Python](https://img.shields.io/pypi/pyversions/termux-vision.svg?style=flat-square)](https://pypi.org/project/termux-vision/)
5
- [![npm](https://img.shields.io/npm/v/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
6
- [![npm downloads](https://img.shields.io/npm/dm/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
7
- [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-vision)
8
- [![Hardware Acceleration](https://img.shields.io/badge/Vulkan-1.1%2B%20Compute-orange?style=flat-square&logo=vulkan)](https://www.vulkan.org/)
9
-
10
- > **Native On-Device Computer Vision & Multimodal Vision-Language Model (VLM) Runtime for Android Termux via Direct Bionic libc & Vulkan Compute Acceleration.**
11
- > *Zero PRoot. Zero Virtualization. 100% Native ARMv8.2-A NEON SIMD & Hardware GPU Offloading.*
12
-
13
- ---
14
-
15
- ## 📑 Table of Contents
16
-
17
- 1. [Overview & Key Capabilities](#1-overview--key-capabilities)
18
- 2. [Installation Guide](#2-installation-guide)
19
- 3. [Enabling Hardware GPU Acceleration (with ameva-runtime)](#3-enabling-hardware-gpu-acceleration-with-ameva-runtime)
20
- 4. [Standardized CLI & Parameter Matrix](#4-standardized-cli--parameter-matrix)
21
- 5. [Dual Engine Code Examples (Python & Node.js)](#5-dual-engine-code-examples-python--nodejs)
22
- 6. [Production Diagnostics (`termux-vision doctor`)](#6-production-diagnostics-termux-vision-doctor)
23
- 7. [Real-World Benchmarks & Hardware Scorecard](#7-real-world-benchmarks--hardware-scorecard)
24
- 8. [Memory & VRAM Architecture (Zero CPU-Mapped VRAM)](#8-memory--vram-architecture)
25
- 9. [Hardware Requirements & Operational Limits](#9-hardware-requirements--operational-limits)
26
- 10. [License & Permissible Use](#10-license--permissible-use)
27
-
28
- ---
29
-
30
- ## 1. Overview & Key Capabilities
31
-
32
- `termux-vision` is an enterprise-grade, on-device multimodal vision inference and spatial computing framework engineered specifically for mobile Android devices. Operating directly against Android's native Bionic libc ABI and host Vulkan compute drivers, `termux-vision` eliminates heavyweight desktop dependencies (OpenCV, TorchVision) and enables high-throughput visual question answering, OCR image captioning, and classical feature extraction directly on edge hardware.
33
-
34
- * **Native Bionic libc ABI Direct Binding**: Runs directly inside Termux user space with zero virtualization indirection, achieving bare-metal compute efficiency.
35
- * **Dual Compute Acceleration**: Integrates ARMv8.2-A DotProd/FP16 SIMD vector instructions with mobile Vulkan compute shader pipelines.
36
- * **Full-Layer GPU Offloading (-ngl 99)**: Dispatches all 99 transformer layers and cross-attention vision projections directly to device GPU VRAM, achieving pure GPU offloading (**0.00 MiB CPU mapped VRAM**).
37
- * **Zero-Dependency Classical Vision Suite**: Native 5-stage Canny edge detector (8-directional BFS hysteresis), Sobel 3x3 filtering, 2D Integral Images, and Haar-like face candidate localization in pure C/Python/JS.
38
- * **Autonomous OOM & LMK Protection**: Enforces atomic model validation (>10MB threshold guard) and quant-pair validation (SmolVLM Q4_K_M + Q8_0 mmproj) to strictly respect Android Low Memory Killer (LMK) bounds.
39
-
40
- ---
41
-
42
- ## 2. Installation Guide
43
-
44
- `termux-vision` is distributed across both Python (PyPI) and Node.js (npm) ecosystems, with official precompiled ARM64 wheel assets published on GitHub Releases.
45
-
46
- ### 2.1 Termux System Prerequisites
47
- Launch Termux and install required native compilers, Vulkan drivers, and image libraries:
48
- ```bash
49
- pkg update -y
50
- pkg install -y python nodejs clang make cmake git termux-api wget vulkan-loader vulkan-headers vulkan-tools opencl-headers python-numpy libjpeg-turbo
51
- ```
52
-
53
- ### 2.2 Python Package Installation
54
-
55
- * **Option A: Install from PyPI (Recommended)**:
56
- ```bash
57
- pip install --upgrade pip setuptools wheel
58
- pip install termux-vision
59
- ```
60
-
61
- * **Option B: Direct GitHub Releases Wheel Asset (SSOT Verified)**:
62
- ```bash
63
- # Download and install the prebuilt v1.4.0 release wheel
64
- pip install https://github.com/uno-km/termux-vision/releases/download/v1.4.0/termux_vision-1.4.0-py3-none-any.whl
65
- ```
66
-
67
- ### 2.3 Node.js / TypeScript CLI Installation
68
- ```bash
69
- # Global CLI installation
70
- npm install -g termux-vision
71
-
72
- # Local project dependency
73
- npm install termux-vision
74
- ```
75
-
76
- ### 2.4 One-Line Bootstrap Installer
77
- Run the universal bootstrap installer to automatically configure repositories, compile native C/C++ acceleration shims, and verify hardware:
78
- ```bash
79
- curl -sL https://raw.githubusercontent.com/uno-km/termux-vision/main/install.sh | bash
80
- ```
81
-
82
- ---
83
-
84
- ## 3. Enabling Hardware GPU Acceleration (with ameva-runtime)
85
-
86
- To unlock mobile GPU acceleration via Vulkan compute shaders and achieve significant speedups over pure CPU execution, install **`termux-vision`** alongside **`ameva-runtime`**:
87
-
88
- ### 🌟 One-Line Installation
89
- ```bash
90
- # Python Environment
91
- pip install termux-vision ameva-runtime termux-llamacpp
92
-
93
- # Node.js Environment
94
- npm install -g termux-vision @ameva/runtime
95
- ```
96
-
97
- ### 🔮 Mobile GPU Silicon Architecture Status
98
-
99
- | GPU Microarchitecture | Silicon / SoC Reference | Status | Optimization Mechanics |
100
- | :--- | :--- | :--- | :--- |
101
- | **Qualcomm Adreno GPU** | Snapdragon 8 Elite (Adreno 830)<br>Snapdragon 8 Gen 1/2/3 (Adreno 730-750) | 🟢 **Production Verified** | Direct Bionic ICD binding, SPIR-V JIT patch (`mul_mat_vec_max_cols = 2`), KGSL Watchdog defense (`GGML_VULKAN_SKIP_CHECKS="999999999"`), micro-batch prefill chunking (`-b 64 -ub 64`). |
102
- | **ARM Mali GPU** | Exynos 2100 (Mali-G78 MP14)<br>Exynos 1380 (Mali-G68 MP5) | 🟢 **Production Verified** | Bionic Vulkan ICD binding, Tile-Based Deferred Rendering (TBDR) memory isolation, MMVQ matrix-vector kernel dispatch (`--tune-mali`). |
103
- | **Samsung Xclipse GPU** | Exynos 2200 / 2400<br>(Xclipse 920 / 940 - AMD RDNA) | 🟡 **In Development (개발 진행 중)** | SPIR-V instruction scheduling and RDNA mobile shader alignment under active engineering. |
104
-
105
- Verify GPU driver detection and hardware readiness:
106
- ```bash
107
- termux-vision doctor
108
- ```
109
-
110
- ---
111
-
112
- ## 4. Standardized CLI & Parameter Matrix
113
-
114
- `termux-vision` strictly complies with the official `uno-km` family CLI standard:
115
-
116
- | Parameter | Alias | Default | Description |
117
- | :--- | :--- | :--- | :--- |
118
- | `-d, --device` | `-b, --backend` | `auto` | Compute acceleration backend: `auto`, `gpu`, `vulkan`, `cpu`, `vulkan-force` |
119
- | `-i, --image` | `--image-path` | *Required* | Path to input image (`.png`, `.jpg`, `.webp`) |
120
- | `-p, --prompt` | *N/A* | `"Describe this image"` | Multimodal text instruction query |
121
- | `-m, --model` | *N/A* | `smolvlm-500m` | GGUF language model path or catalog identifier |
122
- | `--mmproj` | *N/A* | *Auto-paired* | Vision projector GGUF model path (`mmproj-*.gguf`) |
123
- | `-n, --max-tokens` | `--n-predict` | `150` | Maximum number of generated tokens |
124
- | `-c, --ctx-size` | `--ctx` | `2048` | Context window size |
125
- | `-t, --threads` | *N/A* | `auto` | Number of CPU execution threads |
126
- | `-W, --width` | *N/A* | *None* | Explicit image resize width in pixels |
127
- | `-H, --height` | *N/A* | *None* | Explicit image resize height in pixels |
128
- | `--image-size` | *N/A* | *None* | Image resolution preset (e.g. `224x224`, `384x384`) |
129
- | `-q, --quality` | *N/A* | `optimal` | 4-tier resolution preset: `fast` (384px), `optimal` (768px), `high` (1280px), `original` (1:1) |
130
- | `--tune-mali` | *N/A* | `False` | Enable ARM Mali GPU MMVQ tuning (`GGML_VK_FORCE_MMVQ=1`) |
131
- | `--json` | *N/A* | `False` | Emit machine-readable JSON benchmark telemetry |
132
- | `-v, --verbose` | *N/A* | `False` | Print detailed layer offloading and hardware logs |
133
-
134
- ### Practical CLI Usage Examples
135
-
136
- ```bash
137
- # 1. Automated GPU Acceleration (Default auto-routing)
138
- termux-vision vlm photo.jpg -p "What objects are visible in this scene?"
139
-
140
- # 2. Pure GPU Mode on ARM Mali Silicon (Galaxy S21 / A35)
141
- termux-vision vlm photo.jpg -d gpu --tune-mali -p "Describe the text and layout."
142
-
143
- # 3. Pure CPU Fallback Mode (Strict 0 GPU VRAM allocation)
144
- termux-vision vlm photo.jpg -d cpu -t 6 -p "Analyze this diagram."
145
-
146
- # 4. Ultra-Low-Latency Mode (224x224 scaled ViT input)
147
- python tools/vlm_runner.py -i photo.jpg -d gpu --image-size 224x224 -n 60
148
-
149
- # 5. Zero-Dependency Classical Canny Edge Detection
150
- termux-vision canny input.jpg -o edges.png --low 40 --high 120
151
- ```
152
-
153
- ---
154
-
155
- ## 5. Dual Engine Code Examples (Python & Node.js)
156
-
157
- ### 5.1 Python SDK
158
- ```python
159
- import termux_vision as tv
160
-
161
- # 1. Zero-Dependency Classical CV Filtering (sub-10ms execution)
162
- image = tv.io.load_image("document.jpg")
163
- grayscale = tv.transforms.to_grayscale(image)
164
- edges = tv.cv.canny(grayscale, low_threshold=40, high_threshold=120)
165
- tv.io.save_image(edges, "edges.png")
166
-
167
- # 2. On-Device Multimodal VLM Inference (Vulkan GPU Accelerated)
168
- with tv.vlm.load("smolvlm-500m", device="gpu") as engine:
169
- result = engine.describe(
170
- "document.jpg",
171
- prompt="Extract all visible text and summarize key bullet points.",
172
- quality="optimal",
173
- max_tokens=200
174
- )
175
- print(f"Backend: {result.metrics.backend} | TPS: {result.metrics.tokens_per_second:.2f} tok/s")
176
- print(f"Response:
177
- {result.text}")
178
- ```
179
-
180
- ### 5.2 Node.js / TypeScript SDK
181
- ```typescript
182
- import tv from 'termux-vision';
183
-
184
- // 1. Hardware Diagnostic Probe
185
- const doctor = tv.doctor(true);
186
- console.log(`Vulkan GPU: ${doctor.vulkan.status} | Available RAM: ${doctor.hardware.availableRamMb} MB`);
187
-
188
- // 2. Multimodal VLM Inference
189
- const engine = await tv.load({ modelId: 'smolvlm-500m', device: 'gpu' });
190
- const response = await engine.describe('photo.jpg', {
191
- prompt: 'Identify the geometric shapes and colors.',
192
- quality: 'optimal',
193
- maxTokens: 100
194
- });
195
-
196
- console.log(`[${response.metrics.backend.toUpperCase()}] ${response.text}`);
197
- engine.close();
198
- ```
199
-
200
- ---
201
-
202
- ## 6. Production Diagnostics (`termux-vision doctor`)
203
-
204
- Termux-Vision features an integrated hardware diagnostics probe to inspect the host environment before launching inference:
205
-
206
- ```bash
207
- termux-vision doctor
208
- ```
209
-
210
- **Diagnostic Output Profile:**
211
- ```
212
- === termux-vision Diagnostic Doctor ===
213
- Platform : Linux (aarch64) | Android: True
214
- RAM : Total 7812MB | Available 3450MB
215
- CPU Cores: 8 (big.LITTLE Affinity Governor active)
216
- Vulkan : Loader=True | Driver=/system/lib64/libvulkan.so | Status=READY
217
- GPU Soc : ARM Mali-G78 MP14 (Exynos 2100)
218
- Models : 2 installed in ~/.cache/termux-vision/models
219
- Preset : optimal (768px recommended)
220
- ```
221
-
222
- ---
223
-
224
- ## 7. Real-World Benchmarks & Hardware Scorecard
225
-
226
- All metrics represent deterministic ground-truth measurements obtained on physical test devices running Android Termux unrooted, comparing **Moondream2 1.8B f16** and **SmolVLM-500M Instruct (Q4_K_M + Q8_0 mmproj)**.
227
-
228
- | Target Device | SoC & GPU Architecture | Model Architecture | Precision & Weights | Mode | Prompt Processing | Token Generation | Total Latency | Mapped CPU VRAM | Vulkan GPU VRAM | Status / Speedup |
229
- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
230
- | **Samsung Galaxy S25** | Snapdragon 8 Elite<br>Qualcomm Adreno 830 | **Moondream2 1.8B** | Text f16 (2.7GB)<br>+ ViT f16 (868MB) | **GPU (Vulkan 25/25)** | **19.84 tok/s** | **15.00 tok/s** | **39.20 s** | **0.00 MiB** | **2,706.00 MiB** | **Production Verified** |
231
- | **Samsung Galaxy S21 5G** | Exynos 2100<br>ARM Mali-G78 MP14 | SmolVLM-500M-Instruct | Q4_K_M (350MB)<br>+ Q8_0 mmproj (200MB) | **GPU (Vulkan)** | **14.28 tok/s** | **12.65 tok/s** | **21.26 s** | **0.00 MiB** | **1,059.02 MiB** | **+58.9% vs CPU** |
232
- | Samsung Galaxy S21 5G | Exynos 2100<br>8-Core ARMv8.2-A CPU | SmolVLM-500M-Instruct | Q4_K_M (350MB)<br>+ Q8_0 mmproj (200MB) | CPU (NEON) | 8.84 tok/s | 7.96 tok/s | 34.22 s | 1,059.02 MiB | 0.00 MiB | Baseline |
233
- | **Samsung Galaxy A35 5G** | Exynos 1380<br>ARM Mali-G68 MP5 | SmolVLM-500M-Instruct | Q4_K_M (350MB)<br>+ Q8_0 mmproj (200MB) | **GPU (Vulkan)** | **5.67 tok/s** | **5.47 tok/s** | **52.57 s** | **0.00 MiB** | **1,059.02 MiB** | **+55.8% vs CPU** |
234
- | Samsung Galaxy A35 5G | Exynos 1380<br>8-Core ARMv8.2-A CPU | SmolVLM-500M-Instruct | Q4_K_M (350MB)<br>+ Q8_0 mmproj (200MB) | CPU (NEON) | 4.88 tok/s | 3.51 tok/s | 65.34 s | 1,059.02 MiB | 0.00 MiB | Baseline |
235
-
236
- ### 7.1 Empirical Visual Question Answering Verification (Galaxy S25 Adreno 830)
237
-
238
- #### Test Case A: Geometric & Spatial Reasoning (`test_shapes_224x224.png`)
239
- * **Input Image**: Clean canvas with primary geometric primitives (triangles, rectangle).
240
- * **Prompt**: `"Describe the colors and geometric shapes visible in this image."`
241
- * **Model**: `Moondream2 1.8B` (f16 text + f16 ViT mmproj, 2,706 MiB VRAM)
242
- * **Exact Ground-Truth Output**:
243
- > *"The image features a white background with three distinct geometric shapes: two triangles and one rectangle..."*
244
- * **Execution Metrics**:
245
- - GPU Layers Offloaded: **25 / 25 (100% Full GPU)**
246
- - Vulkan VRAM Allocated: **2,706.00 MiB** (CPU Mapped VRAM: **0.00 MiB**)
247
- - Prompt Evaluation: **19.60 tokens/sec** (749 tokens in 38,211 ms)
248
- - Token Generation: **14.97 tokens/sec** (39 tokens in 2,605 ms, 66.80 ms/tok)
249
- - Total Execution Time: **51.24s** (Cold weights loading: 10.4s, CPU ViT projection: 14.1s, GPU decoding: 2.6s)
250
- - Exit Status: `Exit Code 0`
251
-
252
- #### Test Case B: Photorealistic Real-World Scene (`test_elephant_gpu_step2.png`)
253
- * **Input Image**: Diffusion-synthesized high-detail photorealistic scene.
254
- * **Prompt**: `"What animal is this and what is it doing?"`
255
- * **Model**: `Moondream2 1.8B` (f16 text + f16 ViT mmproj)
256
- * **Exact Ground-Truth Output**:
257
- > *"The image shows a large elephant riding on top of a surfboard in the ocean."*
258
- * **Execution Metrics**:
259
- - Prompt Processing: **19.84 tokens/sec** (747 tokens in 37,651 ms)
260
- - Generation Speed: **15.00 tokens/sec** (24 tokens in 1,600 ms, 66.67 ms/tok)
261
- - KGSL Watchdog State: Fully stabilized via `GGML_VULKAN_SKIP_CHECKS="999999999"` (No `ErrorDeviceLost`)
262
- - Exit Status: `Exit Code 0`
263
-
264
- ### 7.2 Key Architectural Discoveries
265
- 1. **Adreno 830 SPIR-V JIT Fix**: Qualcomm's new compiler fails during unrolled vector compilation with `mul_mat_vec_max_cols = 8` (`VK_ERROR_UNKNOWN`). Reducing this parameter to `2` eliminates register spilling and enables full 25/25 layer GPU offloading.
266
- 2. **Android KGSL Watchdog Defense**: Prefilling 729 vision tokens in mobile Vulkan exceeds the 5-second kernel watchdog timer unless micro-batched. Configuring `-b 64 -ub 64` and injecting `GGML_VULKAN_SKIP_CHECKS="999999999"` slices prefill into 1.8s units, preventing `ErrorDeviceLost`.
267
- 3. **Pure GPU Isolation (0.00 MiB CPU VRAM)**: Across both Qualcomm Adreno 830 and ARM Mali (G78/G68), all tensor weights and KV cache reside strictly in Vulkan GPU memory.
268
-
269
- ---
270
-
271
- ## 8. Memory & VRAM Architecture
272
-
273
- ```
274
- +---------------------------------------------------------------+
275
- | Physical Mobile LPDDR4X/LPDDR5 RAM (8 GB) |
276
- +---------------------------------------------------------------+
277
- | |
278
- v v
279
- +-------------------------------+ +-------------------------------+
280
- | Android OS & Framework | | Termux User Space |
281
- | (~3.5 - 4.2 GB) | | (~3.8 - 4.5 GB) |
282
- +-------------------------------+ +-------------------------------+
283
- |
284
- v
285
- +-------------------------------+
286
- | Vulkan Unified Memory |
287
- | - Model Weights: 1059.02 MiB |
288
- | - KV Cache : 384.00 MiB |
289
- | - CPU Mapped : 0.00 MiB |
290
- +-------------------------------+
291
- ```
292
-
293
- ### Quantization & Android LMK Protection
294
- * **SmolVLM-500M (Q4_K_M 350MB + Q8_0 mmproj 200MB)**: Peak working memory stays under ~1.6 GB, well below Android LMK eviction thresholds.
295
- * **Qwen2-VL-2B (Q4_K_M 1.4GB + FP16 mmproj 600MB)**: Requires a minimum of 6GB available RAM; FP32 projector variants (>2.5GB) are automatically rejected to prevent SIGKILL aborts.
296
-
297
- ---
298
-
299
- ## 9. Hardware Requirements & Operational Limits
300
-
301
- | Requirement | Minimum Specification | Recommended Specification |
302
- | :--- | :--- | :--- |
303
- | **Operating System** | Android 10+ (Termux ARM64) | Android 13+ (One UI 5.0+ / Termux Bionic) |
304
- | **Processor (SoC)** | 8-Core ARM64 (Cortex-A55/A76) | Exynos 2100 / Snapdragon 8 Gen 2 or newer |
305
- | **System RAM** | 6 GB LPDDR4X | 8 GB+ LPDDR5 |
306
- | **Vulkan API** | Vulkan 1.1 with SPIR-V Compute | Vulkan 1.2+ with Subgroup 16 arithmetic |
307
- | **Free Storage** | 2.5 GB internal storage | 6.0 GB internal storage |
308
-
309
- ---
310
-
311
- ## 10. License & Permissible Use
312
-
313
- Licensed under the **Apache License, Version 2.0**.
314
- Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)) & AMEVA Open-Source Foundation (AOSF).
315
-
316
- * [Official Documentation Portal](https://uno-km.vercel.app/lib/vision/)
317
- * [AMEVA Open-Source Foundation](https://uno-km.vercel.app/foundation/index.html)
318
- * [GitHub Repository](https://github.com/uno-km/termux-vision)
319
- * [Issue Tracker](https://github.com/uno-km/termux-vision/issues)
1
+ # Termux-Vision: On-Device Computer Vision & Multimodal VLM Framework
2
+
3
+ [![PyPI](https://img.shields.io/pypi/v/termux-vision.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-vision/)
4
+ [![Python](https://img.shields.io/pypi/pyversions/termux-vision.svg?style=flat-square)](https://pypi.org/project/termux-vision/)
5
+ [![npm](https://img.shields.io/npm/v/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
6
+ [![npm downloads](https://img.shields.io/npm/dm/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
7
+ [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-vision)
8
+ [![Hardware Acceleration](https://img.shields.io/badge/Vulkan-1.1%2B%20Compute-orange?style=flat-square&logo=vulkan)](https://www.vulkan.org/)
9
+
10
+ > **Native On-Device Computer Vision & Multimodal Vision-Language Model (VLM) Runtime for Android Termux via Direct Bionic libc & Vulkan Compute Acceleration.**
11
+ > *Zero PRoot. Zero Virtualization. 100% Native ARMv8.2-A NEON SIMD & Hardware GPU Offloading.*
12
+
13
+ ---
14
+
15
+ ## 📑 Table of Contents
16
+
17
+ 1. [Overview & Key Capabilities](#1-overview--key-capabilities)
18
+ 2. [Installation Guide & Prebuilt Installer](#2-installation-guide--prebuilt-installer)
19
+ 3. [Enabling Hardware GPU Acceleration (with ameva-runtime)](#3-enabling-hardware-gpu-acceleration-with-ameva-runtime)
20
+ 4. [Standardized CLI & Parameter Matrix](#4-standardized-cli--parameter-matrix)
21
+ 5. [Dual Engine Code Examples (Python & Node.js)](#5-dual-engine-code-examples-python--nodejs)
22
+ 6. [Production Diagnostics (`termux-vision doctor`)](#6-production-diagnostics-termux-vision-doctor)
23
+ 7. [Real-World Benchmarks & Hardware Scorecard](#7-real-world-benchmarks--hardware-scorecard)
24
+ 8. [Memory & VRAM Architecture (Zero CPU-Mapped VRAM)](#8-memory--vram-architecture)
25
+ 9. [Hardware Requirements & Operational Limits](#9-hardware-requirements--operational-limits)
26
+ 10. [License & Permissible Use](#10-license--permissible-use)
27
+
28
+ ---
29
+
30
+ ## 1. Overview & Key Capabilities
31
+
32
+ `termux-vision` is an enterprise-grade, on-device multimodal vision inference and spatial computing framework engineered specifically for mobile Android devices. Operating directly against Android's native Bionic libc ABI and host Vulkan compute drivers, `termux-vision` eliminates heavyweight desktop dependencies (OpenCV, TorchVision) and enables high-throughput visual question answering, OCR image captioning, and classical feature extraction directly on edge hardware.
33
+
34
+ * **100% Vulkan GPU Compute Canny (`0.23 ms`)**: Chains 3-pass SPIR-V compute shaders (Sobel 3x3, NMS, Hysteresis) entirely within VRAM using `vkCmdPipelineBarrier`, achieving 834x acceleration over Python without CPU memory roundtrips.
35
+ * **Ultra-Fast ARM64 NEON C++ Kernel (`3.02 ms`)**: Permanently eliminates trigonometric `atan2f` via tangent ratio bit quantization and 1-byte direction buffers, running Canny filtering in 3.02ms on Snapdragon 865 and 4.36ms on Exynos 1380.
36
+ * **Prebuilt-Asset-First Idempotent Installer (`0.005s Skip`)**: Automatically provisions verified precompiled ARM64 native binaries in 2 seconds from official releases, guaranteeing zero-build instant skip if assets already exist.
37
+ * **5-Backend Unified CLI Standard**: Enforces `['auto', 'gpu', 'vulkan', 'opencl', 'cpu']` and convenience flags (`--gpu`, `--cpu`, `--opencl`) across all subcommands.
38
+ * **Zero-Deception Fail-Fast Gatekeeper**: Strictly rejects defective text-only binaries lacking `--mmproj` (`E015`) and corrupted weights (`E014`), permanently banning silent fallbacks.
39
+ * **Full-Layer GPU Offloading (-ngl 99)**: Dispatches all transformer layers and cross-attention vision projections directly to device GPU VRAM (**0.00 MiB CPU mapped VRAM**).
40
+
41
+ ---
42
+
43
+ ## 2. Installation Guide & Prebuilt Installer
44
+
45
+ `termux-vision` is distributed across both Python (PyPI) and Node.js (npm) ecosystems, with official precompiled ARM64 wheel assets published on GitHub Releases.
46
+
47
+ ### 2.1 Termux System Prerequisites
48
+ Launch Termux and install required native compilers, Vulkan drivers, and image libraries:
49
+ ```bash
50
+ pkg update -y
51
+ pkg install -y python nodejs clang make cmake git termux-api wget vulkan-loader vulkan-headers vulkan-tools opencl-headers python-numpy libjpeg-turbo
52
+ ```
53
+
54
+ ### 2.2 Python Package Installation
55
+
56
+ * **Option A: Install from PyPI (Recommended)**:
57
+ ```bash
58
+ pip install --upgrade pip setuptools wheel
59
+ pip install termux-vision
60
+ ```
61
+
62
+ * **Option B: Prebuilt Native Engine Provisioning (Idempotent 0.005s)**:
63
+ ```bash
64
+ # Automatically download & unpack verified ARM64 prebuilt assets
65
+ termux-vision install
66
+
67
+ # Optional maintenance flags:
68
+ # termux-vision install --force # Force re-downloading prebuilts
69
+ # termux-vision install --from-source # Force compiling from local C++ source
70
+ # termux-vision install --dry-run # Check integrity without making changes
71
+ ```
72
+
73
+ * **Option C: Direct GitHub Releases Wheel Asset**:
74
+ ```bash
75
+ # Download and install the prebuilt v1.5.0 release wheel
76
+ pip install https://github.com/uno-km/termux-vision/releases/download/v1.5.0/termux_vision-1.5.0-py3-none-any.whl
77
+ ```
78
+
79
+ ### 2.3 Node.js / TypeScript CLI Installation
80
+ ```bash
81
+ # Global CLI installation
82
+ npm install -g termux-vision
83
+
84
+ # Local project dependency
85
+ npm install termux-vision
86
+ ```
87
+
88
+ ### 2.4 One-Line Bootstrap Installer
89
+ Run the universal bootstrap installer to automatically configure repositories, compile native C/C++ acceleration shims, and verify hardware:
90
+ ```bash
91
+ curl -sL https://raw.githubusercontent.com/uno-km/termux-vision/main/install.sh | bash
92
+ ```
93
+
94
+ ---
95
+
96
+ ## 3. Enabling Hardware GPU Acceleration (with ameva-runtime)
97
+
98
+ To unlock mobile GPU acceleration via Vulkan compute shaders and achieve significant speedups over pure CPU execution, install **`termux-vision`** alongside **`ameva-runtime`**:
99
+
100
+ ### 🌟 One-Line Installation
101
+ ```bash
102
+ # Python Environment
103
+ pip install termux-vision ameva-runtime termux-llamacpp
104
+
105
+ # Node.js Environment
106
+ npm install -g termux-vision @ameva/runtime
107
+ ```
108
+
109
+ ### 🔮 Mobile GPU Silicon Architecture Status
110
+
111
+ | GPU Microarchitecture | Silicon / SoC Reference | Status | Optimization Mechanics |
112
+ | :--- | :--- | :--- | :--- |
113
+ | **Qualcomm Adreno GPU** | Snapdragon 8 Elite (Adreno 830)<br>Snapdragon 8 Gen 1/2/3 (Adreno 730-750)<br>Snapdragon 865 (Adreno 650) | 🟢 **Production Verified** | Direct Bionic ICD binding, SPIR-V JIT patch (`mul_mat_vec_max_cols = 2`), KGSL Watchdog defense (`GGML_VULKAN_SKIP_CHECKS="999999999"`), micro-batch prefill chunking (`-b 64 -ub 64`). |
114
+ | **ARM Mali GPU** | Exynos 2100 (Mali-G78 MP14)<br>Exynos 1380 (Mali-G68 MP5) | 🟢 **Production Verified** | Bionic Vulkan ICD binding, Tile-Based Deferred Rendering (TBDR) memory isolation, MMVQ matrix-vector kernel dispatch (`--tune-mali`). |
115
+ | **Samsung Xclipse GPU** | Exynos 2200 / 2400<br>(Xclipse 920 / 940 - AMD RDNA) | 🟡 **In Development (개발 진행 중)** | SPIR-V instruction scheduling and RDNA mobile shader alignment under active engineering. |
116
+
117
+ Verify GPU driver detection and hardware readiness:
118
+ ```bash
119
+ termux-vision doctor
120
+ ```
121
+
122
+ ---
123
+
124
+ ## 4. Standardized CLI & Parameter Matrix
125
+
126
+ `termux-vision` strictly complies with the official `uno-km` family 5-backend CLI standard:
127
+
128
+ | Parameter | Alias | Default | Description |
129
+ | :--- | :--- | :--- | :--- |
130
+ | `-b, --backend` | `-d, --device` | `auto` | Compute acceleration backend: `auto`, `gpu`, `vulkan`, `opencl`, `cpu` |
131
+ | `--gpu` / `--cpu` / `--opencl` | *N/A* | *None* | Convenience shorthand flags for backend routing |
132
+ | `-i, --image` | `--image-path` | *Required* | Path to input image (`.png`, `.jpg`, `.webp`) |
133
+ | `-p, --prompt` | *N/A* | `"Describe this image"` | Multimodal text instruction query |
134
+ | `-m, --model` | *N/A* | `smolvlm-500m` | GGUF language model path or catalog identifier |
135
+ | `--mmproj` | *N/A* | *Auto-paired* | Vision projector GGUF model path (`mmproj-*.gguf`) |
136
+ | `-n, --max-tokens` | `--n-predict` | `150` | Maximum number of generated tokens |
137
+ | `-c, --ctx-size` | `--ctx` | `2048` | Context window size |
138
+ | `-t, --threads` | *N/A* | `auto` | Number of CPU execution threads |
139
+ | `--image-size` | *N/A* | *None* | Image resolution preset (e.g. `224x224`, `384x384`) |
140
+ | `-q, --quality` | *N/A* | `optimal` | 4-tier resolution preset: `fast` (384px), `optimal` (768px), `high` (1280px), `original` (1:1) |
141
+ | `--tune-mali` | *N/A* | `False` | Enable ARM Mali GPU MMVQ tuning (`GGML_VK_FORCE_MMVQ=1`) |
142
+ | `--json` | *N/A* | `False` | Emit machine-readable JSON benchmark telemetry |
143
+ | `-v, --verbose` | *N/A* | `False` | Print detailed layer offloading and hardware logs |
144
+
145
+ ### Practical CLI Usage Examples
146
+
147
+ ```bash
148
+ # 1. 100% Vulkan GPU Canny Edge Detection (0.23 ms on Adreno 830)
149
+ termux-vision canny photo.jpg -o edges.png --gpu --low 40 --high 120
150
+
151
+ # 2. Ultra-Fast NEON C++ Canny Edge Detection (3.02 ms on S20 CPU)
152
+ termux-vision canny photo.jpg -o edges.png --cpu
153
+
154
+ # 3. Multimodal VLM Inference with Automated GPU Routing
155
+ termux-vision vlm photo.jpg -p "What objects are visible in this scene?"
156
+
157
+ # 4. Pure GPU Mode on ARM Mali Silicon (Galaxy S21 / A35)
158
+ termux-vision vlm photo.jpg -d gpu --tune-mali -p "Describe the text and layout."
159
+
160
+ # 5. Prebuilt Native Binary Provisioning (0.005s Idempotent Skip)
161
+ termux-vision install
162
+ ```
163
+
164
+ ---
165
+
166
+ ## 5. Dual Engine Code Examples (Python & Node.js)
167
+
168
+ ### 5.1 Python SDK
169
+ ```python
170
+ import termux_vision as tv
171
+
172
+ # 1. Hardware-Accelerated Canny Edge Detection (0.23ms Vulkan GPU / 3.02ms NEON CPU)
173
+ image = tv.io.load_image("document.jpg")
174
+ grayscale = tv.transforms.to_grayscale(image)
175
+ edges = tv.cv.canny(grayscale, low_threshold=40, high_threshold=120, backend="auto")
176
+ tv.io.save_image(edges, "edges.png")
177
+
178
+ # 2. On-Device Multimodal VLM Inference (Vulkan GPU Accelerated)
179
+ with tv.vlm.load("smolvlm-500m", device="gpu") as engine:
180
+ result = engine.describe(
181
+ "document.jpg",
182
+ prompt="Extract all visible text and summarize key bullet points.",
183
+ quality="optimal",
184
+ max_tokens=200
185
+ )
186
+ print(f"Backend: {result.metrics.backend} | TPS: {result.metrics.tokens_per_second:.2f} tok/s")
187
+ print(f"Response:\n{result.text}")
188
+ ```
189
+
190
+ ### 5.2 Node.js / TypeScript SDK
191
+ ```typescript
192
+ import tv from 'termux-vision';
193
+
194
+ // 1. Hardware Diagnostic Probe
195
+ const doctor = tv.doctor(true);
196
+ console.log(`Vulkan GPU: ${doctor.vulkan.status} | Available RAM: ${doctor.hardware.availableRamMb} MB`);
197
+
198
+ // 2. Multimodal VLM Inference
199
+ const engine = await tv.load({ modelId: 'smolvlm-500m', device: 'gpu' });
200
+ const response = await engine.describe('photo.jpg', {
201
+ prompt: 'Identify the geometric shapes and colors.',
202
+ quality: 'optimal',
203
+ maxTokens: 100
204
+ });
205
+
206
+ console.log(`[${response.metrics.backend.toUpperCase()}] ${response.text}`);
207
+ engine.close();
208
+ ```
209
+
210
+ ---
211
+
212
+ ## 6. Production Diagnostics (`termux-vision doctor`)
213
+
214
+ Termux-Vision features an integrated hardware diagnostics probe to inspect the host environment before launching inference:
215
+
216
+ ```bash
217
+ termux-vision doctor
218
+ ```
219
+
220
+ **Diagnostic Output Profile:**
221
+ ```
222
+ === termux-vision Diagnostic Doctor ===
223
+ Platform : Linux (aarch64) | Android: True
224
+ RAM : Total 7812MB | Available 3450MB
225
+ CPU Cores: 8 (big.LITTLE Affinity Governor active)
226
+ Vulkan : Loader=True | Driver=/system/lib64/libvulkan.so | Status=READY
227
+ GPU Soc : ARM Mali-G78 MP14 (Exynos 2100)
228
+ Models : 2 installed in ~/.cache/termux-vision/models
229
+ Preset : optimal (768px recommended)
230
+ ```
231
+
232
+ ---
233
+
234
+ ## 7. Real-World Benchmarks & Hardware Scorecard
235
+
236
+ ### 7.1 Classical Vision Filtering Latency (512x512 Image, Physical Devices)
237
+
238
+ | Algorithm / Kernel | Architecture / Acceleration | Execution Latency | Memory Overhead | Status / Verified Device |
239
+ | :--- | :--- | :--- | :--- | :--- |
240
+ | **100% Vulkan GPU Compute Canny** | **3-Pass SPIR-V Compute VRAM Chain** | **0.23 ms (Min 0.18 ms)** | **0.00 MiB CPU VRAM** | **Production (S25 Adreno 830, 834x Speedup)** |
241
+ | **ARM64 NEON C++ Canny Engine** | **Tangent-Ratio Bit Quantization (No atan2f)** | **3.02 ms ~ 3.42 ms** | **1.0 MB (uint8 buffer)** | **Production (S20: 3.02ms, S25: 3.42ms)** |
242
+ | ARM64 NEON C++ Canny Engine | Exynos 1380 Cortex-A78 NEON | 4.36 ms | 1.0 MB | Production (Galaxy A35) |
243
+ | Sobel 3x3 Gradient Convolution | ARM NEON Vectorized | 1.1 ms | 0.5 MB | Production |
244
+ | Gaussian Blur 5x5 Kernel | Separable 1D Conv | 1.8 ms | 0.5 MB | Production |
245
+ | 2D Integral Image (SAT) | Row/Col Prefix Sum | 1.2 ms | 2.0 MB | Production |
246
+ | Haar Cascade Face Detection | Candidate Classifier | 12.5 ms | 2.2 MB | Production |
247
+
248
+ ### 7.2 On-Device Multimodal VLM Benchmark
249
+
250
+ | Target Device | SoC & GPU Architecture | Model Architecture | Mode | Prompt Processing | Token Generation | Mapped CPU VRAM | Vulkan GPU VRAM | Status / Speedup |
251
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
252
+ | **Samsung Galaxy S25** | Snapdragon 8 Elite<br>Adreno 830 | **Moondream2 1.8B f16** | **GPU (Vulkan 25/25)** | **19.84 tok/s** | **15.00 tok/s** | **0.00 MiB** | **2,706.00 MiB** | **Production Verified** |
253
+ | **Samsung Galaxy S21 5G** | Exynos 2100<br>Mali-G78 MP14 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **14.28 tok/s** | **12.65 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+58.9% vs CPU** |
254
+ | Samsung Galaxy S21 5G | Exynos 2100<br>8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 8.84 tok/s | 7.96 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
255
+ | **Samsung Galaxy A35 5G** | Exynos 1380<br>Mali-G68 MP5 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **5.67 tok/s** | **5.47 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+55.8% vs CPU** |
256
+ | Samsung Galaxy A35 5G | Exynos 1380<br>8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 4.88 tok/s | 3.51 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
257
+
258
+ ---
259
+
260
+ ## 8. Memory & VRAM Architecture
261
+
262
+ ```
263
+ +---------------------------------------------------------------+
264
+ | Physical Mobile LPDDR4X/LPDDR5 RAM (8 GB) |
265
+ +---------------------------------------------------------------+
266
+ | |
267
+ v v
268
+ +-------------------------------+ +-------------------------------+
269
+ | Android OS & Framework | | Termux User Space |
270
+ | (~3.5 - 4.2 GB) | | (~3.8 - 4.5 GB) |
271
+ +-------------------------------+ +-------------------------------+
272
+ |
273
+ v
274
+ +-------------------------------+
275
+ | Vulkan Unified Memory |
276
+ | - Model Weights: 1059.02 MiB |
277
+ | - KV Cache : 384.00 MiB |
278
+ | - CPU Mapped : 0.00 MiB |
279
+ +-------------------------------+
280
+ ```
281
+
282
+ ---
283
+
284
+ ## 9. Hardware Requirements & Operational Limits
285
+
286
+ | Requirement | Minimum Specification | Recommended Specification |
287
+ | :--- | :--- | :--- |
288
+ | **Operating System** | Android 10+ (Termux ARM64) | Android 13+ (One UI 5.0+ / Termux Bionic) |
289
+ | **Processor (SoC)** | 8-Core ARM64 (Cortex-A55/A76) | Exynos 2100 / Snapdragon 8 Gen 2 or newer |
290
+ | **System RAM** | 6 GB LPDDR4X | 8 GB+ LPDDR5 |
291
+ | **Vulkan API** | Vulkan 1.1 with SPIR-V Compute | Vulkan 1.2+ with Subgroup 16 arithmetic |
292
+ | **Free Storage** | 2.5 GB internal storage | 6.0 GB internal storage |
293
+
294
+ ---
295
+
296
+ ## 10. License & Permissible Use
297
+
298
+ Licensed under the **Apache License, Version 2.0**.
299
+ Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)) & AMEVA Open-Source Foundation (AOSF).
300
+
301
+ * [Official Documentation Portal](https://uno-km.vercel.app/lib/vision/)
302
+ * [AMEVA Open-Source Foundation](https://uno-km.vercel.app/foundation/index.html)
303
+ * [GitHub Repository](https://github.com/uno-km/termux-vision)
304
+ * [Issue Tracker](https://github.com/uno-km/termux-vision/issues)