termux-vision 1.4.4 → 1.4.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (5) hide show
  1. package/README.md +319 -319
  2. package/README.pypi.md +121 -121
  3. package/bin/cli.js +51 -51
  4. package/lib/vlm.js +349 -349
  5. package/package.json +55 -55
package/README.pypi.md CHANGED
@@ -1,121 +1,121 @@
1
- # Termux-Vision: On-Device Computer Vision & Multimodal VLM Framework
2
-
3
- [![PyPI](https://img.shields.io/pypi/v/termux-vision.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-vision/)
4
- [![Python](https://img.shields.io/pypi/pyversions/termux-vision.svg?style=flat-square)](https://pypi.org/project/termux-vision/)
5
- [![npm](https://img.shields.io/npm/v/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
6
- [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-vision)
7
-
8
- > **Native On-Device Computer Vision & Multimodal Vision-Language Model (VLM) Runtime for Android Termux via Direct Bionic libc & Vulkan Compute Acceleration.**
9
- > *Zero PRoot. Zero Virtualization. 100% Native ARMv8.2-A NEON SIMD & Hardware GPU Offloading.*
10
-
11
- ---
12
-
13
- ## 1. Overview & Key Capabilities
14
-
15
- `termux-vision` is an enterprise-grade, on-device multimodal vision inference and spatial computing framework engineered specifically for mobile Android devices. Operating directly against Android's native Bionic libc ABI and host Vulkan compute drivers, `termux-vision` eliminates heavyweight desktop dependencies (OpenCV, TorchVision) and enables high-throughput visual question answering, OCR image captioning, and classical feature extraction directly on edge hardware.
16
-
17
- * **Native Bionic libc ABI Direct Binding**: Runs directly inside Termux user space with zero virtualization indirection.
18
- * **Dual Compute Acceleration**: Integrates ARMv8.2-A DotProd/FP16 SIMD vector instructions with mobile Vulkan compute shader pipelines.
19
- * **Full-Layer GPU Offloading (-ngl 99)**: Dispatches all 99 transformer layers and cross-attention vision projections directly to device GPU VRAM (**0.00 MiB CPU mapped VRAM**).
20
- * **Zero-Dependency Classical Vision Suite**: Native 5-stage Canny edge detector (8-directional BFS hysteresis), Sobel 3x3 filtering, 2D Integral Images, and Haar-like face candidate localization.
21
-
22
- ---
23
-
24
- ## 2. Installation Guide
25
-
26
- ### 2.1 Python Package Installation (PyPI)
27
- ```bash
28
- pip install --upgrade pip setuptools wheel
29
- pip install termux-vision
30
- ```
31
-
32
- ### 2.2 Direct GitHub Releases Wheel Asset
33
- ```bash
34
- pip install https://github.com/uno-km/termux-vision/releases/download/v1.4.0/termux_vision-1.4.0-py3-none-any.whl
35
- ```
36
-
37
- ### 2.3 One-Line Bootstrap Installer
38
- ```bash
39
- curl -sL https://raw.githubusercontent.com/uno-km/termux-vision/main/install.sh | bash
40
- ```
41
-
42
- ---
43
-
44
- ## 3. GPU Hardware Acceleration (`ameva-runtime`)
45
-
46
- ```bash
47
- pip install termux-vision ameva-runtime termux-llamacpp
48
- ```
49
-
50
- ### Silicon Architecture Support Status
51
- * **Qualcomm Adreno GPU (Snapdragon 8 Elite / Adreno 830, Adreno 7xx)**: Production Verified & Supported (Full 25/25 layer GPU offloading, 15.00 tok/s on Moondream2 1.8B f16, SPIR-V JIT patch, KGSL watchdog defense via `GGML_VULKAN_SKIP_CHECKS="999999999"`).
52
- * **ARM Mali GPU (Mali-G78, Mali-G68, etc.)**: Production Verified & Supported (Pure GPU offloading, 0.00 MiB CPU Mapped VRAM, MMVQ tuning via `--tune-mali`).
53
- * **Samsung Xclipse GPU (Xclipse 920 / 940 - AMD RDNA)**: Under Active Engineering (In Progress / 개발 진행 중).
54
-
55
- Run hardware diagnostics:
56
- ```bash
57
- termux-vision doctor
58
- ```
59
-
60
- ---
61
-
62
- ## 4. Standardized CLI & Parameter Matrix
63
-
64
- | Parameter | Alias | Default | Description |
65
- | :--- | :--- | :--- | :--- |
66
- | `-d, --device` | `-b, --backend` | `auto` | Compute acceleration backend: `auto`, `gpu`, `vulkan`, `cpu`, `vulkan-force` |
67
- | `-i, --image` | `--image-path` | *Required* | Path to input image (`.png`, `.jpg`, `.webp`) |
68
- | `-p, --prompt` | *N/A* | `"Describe this image"` | Multimodal text instruction query |
69
- | `-m, --model` | *N/A* | `smolvlm-500m` | GGUF language model path or catalog identifier |
70
- | `--mmproj` | *N/A* | *Auto-paired* | Vision projector GGUF model path (`mmproj-*.gguf`) |
71
- | `-n, --max-tokens` | `--n-predict` | `150` | Maximum number of generated tokens |
72
- | `-c, --ctx-size` | `--ctx` | `2048` | Context window size |
73
- | `-t, --threads` | *N/A* | `auto` | Number of CPU execution threads |
74
- | `-q, --quality` | *N/A* | `optimal` | 4-tier resolution preset: `fast` (384px), `optimal` (768px), `high` (1280px), `original` (1:1) |
75
- | `--tune-mali` | *N/A* | `False` | Enable ARM Mali GPU MMVQ tuning (`GGML_VK_FORCE_MMVQ=1`) |
76
- | `--json` | *N/A* | `False` | Emit machine-readable JSON benchmark telemetry |
77
-
78
- ### CLI Example
79
- ```bash
80
- # GPU-Accelerated Multimodal VLM Inference
81
- termux-vision vlm photo.jpg -d gpu --tune-mali -p "What objects are visible?"
82
-
83
- # Classical 5-stage Canny Edge Detection
84
- termux-vision canny photo.jpg -o edges.png --low 40 --high 120
85
- ```
86
-
87
- ---
88
-
89
- ## 5. Python SDK Quickstart
90
-
91
- ```python
92
- import termux_vision as tv
93
-
94
- # 1. Zero-Dependency Classical CV Filtering
95
- image = tv.io.load_image("document.jpg")
96
- grayscale = tv.transforms.to_grayscale(image)
97
- edges = tv.cv.canny(grayscale, low_threshold=40, high_threshold=120)
98
- tv.io.save_image(edges, "edges.png")
99
-
100
- # 2. On-Device Multimodal VLM Inference
101
- with tv.vlm.load("smolvlm-500m", device="gpu") as engine:
102
- result = engine.describe("document.jpg", prompt="Summarize this document.", quality="optimal")
103
- print(f"TPS: {result.metrics.tokens_per_second:.2f} tok/s | Output: {result.text}")
104
- ```
105
-
106
- ---
107
-
108
- ## 6. Real-World Benchmarks & Hardware Scorecard
109
-
110
- | Target Device | SoC & GPU | Model Architecture | Mode | Prompt Processing | Token Generation | Mapped CPU VRAM | Vulkan GPU VRAM | Status / Speedup |
111
- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
112
- | **Samsung Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | **Moondream2 1.8B f16** | **GPU (Vulkan 25/25)** | **19.84 tok/s** | **15.00 tok/s** | **0.00 MiB** | **2,706.00 MiB** | **Verified (Full GPU)** |
113
- | **Samsung Galaxy S21 5G** | Exynos 2100 / Mali-G78 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **14.28 tok/s** | **12.65 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+58.9% vs CPU** |
114
- | Samsung Galaxy S21 5G | Exynos 2100 / 8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 8.84 tok/s | 7.96 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
115
- | **Samsung Galaxy A35 5G** | Exynos 1380 / Mali-G68 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **5.67 tok/s** | **5.47 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+55.8% vs CPU** |
116
- | Samsung Galaxy A35 5G | Exynos 1380 / 8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 4.88 tok/s | 3.51 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
117
-
118
- ---
119
-
120
- ## 7. License
121
- Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)) & AOSF.
1
+ # Termux-Vision: On-Device Computer Vision & Multimodal VLM Framework
2
+
3
+ [![PyPI](https://img.shields.io/pypi/v/termux-vision.svg?style=flat-square&color=0369a1)](https://pypi.org/project/termux-vision/)
4
+ [![Python](https://img.shields.io/pypi/pyversions/termux-vision.svg?style=flat-square)](https://pypi.org/project/termux-vision/)
5
+ [![npm](https://img.shields.io/npm/v/termux-vision.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/termux-vision)
6
+ [![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/termux-vision)
7
+
8
+ > **Native On-Device Computer Vision & Multimodal Vision-Language Model (VLM) Runtime for Android Termux via Direct Bionic libc & Vulkan Compute Acceleration.**
9
+ > *Zero PRoot. Zero Virtualization. 100% Native ARMv8.2-A NEON SIMD & Hardware GPU Offloading.*
10
+
11
+ ---
12
+
13
+ ## 1. Overview & Key Capabilities
14
+
15
+ `termux-vision` is an enterprise-grade, on-device multimodal vision inference and spatial computing framework engineered specifically for mobile Android devices. Operating directly against Android's native Bionic libc ABI and host Vulkan compute drivers, `termux-vision` eliminates heavyweight desktop dependencies (OpenCV, TorchVision) and enables high-throughput visual question answering, OCR image captioning, and classical feature extraction directly on edge hardware.
16
+
17
+ * **Native Bionic libc ABI Direct Binding**: Runs directly inside Termux user space with zero virtualization indirection.
18
+ * **Dual Compute Acceleration**: Integrates ARMv8.2-A DotProd/FP16 SIMD vector instructions with mobile Vulkan compute shader pipelines.
19
+ * **Full-Layer GPU Offloading (-ngl 99)**: Dispatches all 99 transformer layers and cross-attention vision projections directly to device GPU VRAM (**0.00 MiB CPU mapped VRAM**).
20
+ * **Zero-Dependency Classical Vision Suite**: Native 5-stage Canny edge detector (8-directional BFS hysteresis), Sobel 3x3 filtering, 2D Integral Images, and Haar-like face candidate localization.
21
+
22
+ ---
23
+
24
+ ## 2. Installation Guide
25
+
26
+ ### 2.1 Python Package Installation (PyPI)
27
+ ```bash
28
+ pip install --upgrade pip setuptools wheel
29
+ pip install termux-vision
30
+ ```
31
+
32
+ ### 2.2 Direct GitHub Releases Wheel Asset
33
+ ```bash
34
+ pip install https://github.com/uno-km/termux-vision/releases/download/v1.4.0/termux_vision-1.4.0-py3-none-any.whl
35
+ ```
36
+
37
+ ### 2.3 One-Line Bootstrap Installer
38
+ ```bash
39
+ curl -sL https://raw.githubusercontent.com/uno-km/termux-vision/main/install.sh | bash
40
+ ```
41
+
42
+ ---
43
+
44
+ ## 3. GPU Hardware Acceleration (`ameva-runtime`)
45
+
46
+ ```bash
47
+ pip install termux-vision ameva-runtime termux-llamacpp
48
+ ```
49
+
50
+ ### Silicon Architecture Support Status
51
+ * **Qualcomm Adreno GPU (Snapdragon 8 Elite / Adreno 830, Adreno 7xx)**: Production Verified & Supported (Full 25/25 layer GPU offloading, 15.00 tok/s on Moondream2 1.8B f16, SPIR-V JIT patch, KGSL watchdog defense via `GGML_VULKAN_SKIP_CHECKS="999999999"`).
52
+ * **ARM Mali GPU (Mali-G78, Mali-G68, etc.)**: Production Verified & Supported (Pure GPU offloading, 0.00 MiB CPU Mapped VRAM, MMVQ tuning via `--tune-mali`).
53
+ * **Samsung Xclipse GPU (Xclipse 920 / 940 - AMD RDNA)**: Under Active Engineering (In Progress / 개발 진행 중).
54
+
55
+ Run hardware diagnostics:
56
+ ```bash
57
+ termux-vision doctor
58
+ ```
59
+
60
+ ---
61
+
62
+ ## 4. Standardized CLI & Parameter Matrix
63
+
64
+ | Parameter | Alias | Default | Description |
65
+ | :--- | :--- | :--- | :--- |
66
+ | `-d, --device` | `-b, --backend` | `auto` | Compute acceleration backend: `auto`, `gpu`, `vulkan`, `cpu`, `vulkan-force` |
67
+ | `-i, --image` | `--image-path` | *Required* | Path to input image (`.png`, `.jpg`, `.webp`) |
68
+ | `-p, --prompt` | *N/A* | `"Describe this image"` | Multimodal text instruction query |
69
+ | `-m, --model` | *N/A* | `smolvlm-500m` | GGUF language model path or catalog identifier |
70
+ | `--mmproj` | *N/A* | *Auto-paired* | Vision projector GGUF model path (`mmproj-*.gguf`) |
71
+ | `-n, --max-tokens` | `--n-predict` | `150` | Maximum number of generated tokens |
72
+ | `-c, --ctx-size` | `--ctx` | `2048` | Context window size |
73
+ | `-t, --threads` | *N/A* | `auto` | Number of CPU execution threads |
74
+ | `-q, --quality` | *N/A* | `optimal` | 4-tier resolution preset: `fast` (384px), `optimal` (768px), `high` (1280px), `original` (1:1) |
75
+ | `--tune-mali` | *N/A* | `False` | Enable ARM Mali GPU MMVQ tuning (`GGML_VK_FORCE_MMVQ=1`) |
76
+ | `--json` | *N/A* | `False` | Emit machine-readable JSON benchmark telemetry |
77
+
78
+ ### CLI Example
79
+ ```bash
80
+ # GPU-Accelerated Multimodal VLM Inference
81
+ termux-vision vlm photo.jpg -d gpu --tune-mali -p "What objects are visible?"
82
+
83
+ # Classical 5-stage Canny Edge Detection
84
+ termux-vision canny photo.jpg -o edges.png --low 40 --high 120
85
+ ```
86
+
87
+ ---
88
+
89
+ ## 5. Python SDK Quickstart
90
+
91
+ ```python
92
+ import termux_vision as tv
93
+
94
+ # 1. Zero-Dependency Classical CV Filtering
95
+ image = tv.io.load_image("document.jpg")
96
+ grayscale = tv.transforms.to_grayscale(image)
97
+ edges = tv.cv.canny(grayscale, low_threshold=40, high_threshold=120)
98
+ tv.io.save_image(edges, "edges.png")
99
+
100
+ # 2. On-Device Multimodal VLM Inference
101
+ with tv.vlm.load("smolvlm-500m", device="gpu") as engine:
102
+ result = engine.describe("document.jpg", prompt="Summarize this document.", quality="optimal")
103
+ print(f"TPS: {result.metrics.tokens_per_second:.2f} tok/s | Output: {result.text}")
104
+ ```
105
+
106
+ ---
107
+
108
+ ## 6. Real-World Benchmarks & Hardware Scorecard
109
+
110
+ | Target Device | SoC & GPU | Model Architecture | Mode | Prompt Processing | Token Generation | Mapped CPU VRAM | Vulkan GPU VRAM | Status / Speedup |
111
+ | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
112
+ | **Samsung Galaxy S25** | Snapdragon 8 Elite / Adreno 830 | **Moondream2 1.8B f16** | **GPU (Vulkan 25/25)** | **19.84 tok/s** | **15.00 tok/s** | **0.00 MiB** | **2,706.00 MiB** | **Verified (Full GPU)** |
113
+ | **Samsung Galaxy S21 5G** | Exynos 2100 / Mali-G78 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **14.28 tok/s** | **12.65 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+58.9% vs CPU** |
114
+ | Samsung Galaxy S21 5G | Exynos 2100 / 8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 8.84 tok/s | 7.96 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
115
+ | **Samsung Galaxy A35 5G** | Exynos 1380 / Mali-G68 | SmolVLM-500M-Instruct | **GPU (Vulkan)** | **5.67 tok/s** | **5.47 tok/s** | **0.00 MiB** | **1,059.02 MiB** | **+55.8% vs CPU** |
116
+ | Samsung Galaxy A35 5G | Exynos 1380 / 8-Core CPU | SmolVLM-500M-Instruct | CPU (NEON) | 4.88 tok/s | 3.51 tok/s | 1,059.02 MiB | 0.00 MiB | Baseline |
117
+
118
+ ---
119
+
120
+ ## 7. License
121
+ Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)) & AOSF.
package/bin/cli.js CHANGED
@@ -1,51 +1,51 @@
1
- #!/usr/bin/env node
2
- /**
3
- * AMEVA Standard Node.js CLI Runner for termux_vision.
4
- * Automatically resolves Python 3 environment and dispatches to python -m termux_vision.
5
- */
6
- const fs = require('fs');
7
- const { spawn, execSync } = require('child_process');
8
-
9
- function findPython() {
10
- if (process.env.PYTHON && fs.existsSync(process.env.PYTHON)) {
11
- return process.env.PYTHON;
12
- }
13
- const termuxBin = '/data/data/com.termux/files/usr/bin/python3';
14
- if (fs.existsSync(termuxBin)) {
15
- return termuxBin;
16
- }
17
- const termuxBinAlt = '/data/data/com.termux/files/usr/bin/python';
18
- if (fs.existsSync(termuxBinAlt)) {
19
- return termuxBinAlt;
20
- }
21
- const candidates = ['python3', 'python'];
22
- for (const cmd of candidates) {
23
- try {
24
- const checkCmd = process.platform === 'win32' ? `where ${cmd}` : `command -v ${cmd}`;
25
- const res = execSync(checkCmd, { stdio: ['ignore', 'pipe', 'ignore'] }).toString().trim();
26
- if (res) return cmd;
27
- } catch (_) {}
28
- }
29
- return 'python3';
30
- }
31
-
32
- const pythonBin = findPython();
33
- const args = ['-m', 'termux_vision', ...process.argv.slice(2)];
34
-
35
- const child = spawn(pythonBin, args, {
36
- stdio: 'inherit',
37
- env: process.env
38
- });
39
-
40
- child.on('error', (err) => {
41
- console.error(`[${'termux_vision'}] Failed to spawn python process (${pythonBin}):`, err.message);
42
- process.exit(1);
43
- });
44
-
45
- child.on('exit', (code, signal) => {
46
- if (signal) {
47
- process.kill(process.pid, signal);
48
- } else {
49
- process.exit(code || 0);
50
- }
51
- });
1
+ #!/usr/bin/env node
2
+ /**
3
+ * AMEVA Standard Node.js CLI Runner for termux_vision.
4
+ * Automatically resolves Python 3 environment and dispatches to python -m termux_vision.
5
+ */
6
+ const fs = require('fs');
7
+ const { spawn, execSync } = require('child_process');
8
+
9
+ function findPython() {
10
+ if (process.env.PYTHON && fs.existsSync(process.env.PYTHON)) {
11
+ return process.env.PYTHON;
12
+ }
13
+ const termuxBin = '/data/data/com.termux/files/usr/bin/python3';
14
+ if (fs.existsSync(termuxBin)) {
15
+ return termuxBin;
16
+ }
17
+ const termuxBinAlt = '/data/data/com.termux/files/usr/bin/python';
18
+ if (fs.existsSync(termuxBinAlt)) {
19
+ return termuxBinAlt;
20
+ }
21
+ const candidates = ['python3', 'python'];
22
+ for (const cmd of candidates) {
23
+ try {
24
+ const checkCmd = process.platform === 'win32' ? `where ${cmd}` : `command -v ${cmd}`;
25
+ const res = execSync(checkCmd, { stdio: ['ignore', 'pipe', 'ignore'] }).toString().trim();
26
+ if (res) return cmd;
27
+ } catch (_) {}
28
+ }
29
+ return 'python3';
30
+ }
31
+
32
+ const pythonBin = findPython();
33
+ const args = ['-m', 'termux_vision', ...process.argv.slice(2)];
34
+
35
+ const child = spawn(pythonBin, args, {
36
+ stdio: 'inherit',
37
+ env: process.env
38
+ });
39
+
40
+ child.on('error', (err) => {
41
+ console.error(`[${'termux_vision'}] Failed to spawn python process (${pythonBin}):`, err.message);
42
+ process.exit(1);
43
+ });
44
+
45
+ child.on('exit', (code, signal) => {
46
+ if (signal) {
47
+ process.kill(process.pid, signal);
48
+ } else {
49
+ process.exit(code || 0);
50
+ }
51
+ });