@bismawy/pi-vision-watcher 1.0.8 → 1.0.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +81 -84
  2. package/package.json +2 -2
package/README.md CHANGED
@@ -1,80 +1,67 @@
1
- # đŸ‘ī¸ pi-vision-watcher
1
+ # đŸ‘ī¸ @bismawy/pi-vision-watcher
2
2
 
3
- **Give text-only [pi](https://github.com/earendil-works/pi-coding-agent) models vision**
3
+ **Give text-only [pi](https://github.com/earendil-works/pi-coding-agent) models vision capabilities.**
4
4
 
5
- Describe images using an authenticated vision model of your choice, then seamlessly hand off the text descriptions to non-vision models.
5
+ Seamlessly inspect, describe, and convert visual inputs (screenshots, mockups, terminal errors, clipboard pastes) into structured descriptions using your preferred vision model, and hand them off to text-only coding models without interrupting your workflow.
6
6
 
7
7
  [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
8
8
  [![npm](https://img.shields.io/npm/v/@bismawy/pi-vision-watcher)](https://www.npmjs.com/package/@bismawy/pi-vision-watcher)
9
9
  [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
10
10
 
11
- ---
12
-
13
- ## The Problem
14
-
15
- Some of the best coding models are text-only. When you attach a screenshot, diagram, or UI mock, they either ignore it or fail the request entirely. Switching models just to read an image interrupts your workflow.
16
-
17
- ## The Solution
18
-
19
- `pi-vision-watcher` bridges this gap automatically:
20
- - **Interactive Picker:** Pick any vision model from your authenticated providers with `/vision-watcher`.
21
- - **Automatic Handoff:** Whenever a non-vision model receives an image (via paste, attachment, or the `read` tool), the image is described behind the scenes and swapped for rich descriptive text before reaching the model.
22
- - **Batched & Cached:** Uses a DataLoader pattern so multiple images in a turn coalesce into a **single batched vision request**, cached by SHA-256 hash.
11
+ ![pi-vision-watcher](./assets/screenshot.webp)
23
12
 
24
13
  ---
25
14
 
26
- ## ✨ Features
27
-
28
- - đŸŽ¯ **Connected-Only Model Picker** — `/vision-watcher` filters out unconfigured providers, showing only models you actually have credentials for (`/login`, `models.json`, or environment variables). Vision-capable models are highlighted with đŸ‘ī¸.
29
- - ⚡ **DataLoader Batching** — Multiple images from parallel `read` calls or multi-image attachments merge into ONE batched vision call during the tool-result phase, eliminating latency bottlenecks.
30
- - 🧠 **Thinking & Reasoning Support** — Configure reasoning effort (`/vision-watcher thinking <level>`) for reasoning-capable vision models (e.g. OpenAI o-series, Claude, DeepSeek).
31
- - 🔄 **Fallback Chains** — Specify backup vision models that automatically take over if your primary describer is unavailable or encounters rate limits.
32
- - 🚀 **Paste-Time Prewarm (Opt-in)** — Describe pasted images the moment the path lands in the prompt editor before you even press Enter.
33
- - đŸ“Ŧ **Async Clipboard Fallback (Opt-in)** — Races direct reads against asynchronous description delivery to prevent stalling.
34
- - 💾 **LRU Hash Caching** — Prevents duplicate calls for identical images across conversation turns.
35
- - đŸ›Ąī¸ **Graceful Degradation** — Never crashes your agent turn. If a description fails, a clean `[Image: description unavailable]` placeholder is provided and logged to `~/.pi/agent/logs/pi-vision-watcher/errors.log`.
36
-
37
- ---
38
-
39
- ## đŸ“Ļ Install
15
+ ## ⚡ Quick Start
40
16
 
17
+ ### 1. Installation
41
18
  ```bash
42
19
  pi install npm:@bismawy/pi-vision-watcher
43
20
  ```
21
+ *(Or install directly from Git: `pi install git:github.com/bismawy/pi-vision-watcher`)*
44
22
 
45
- Alternatively, install directly from GitHub:
46
-
23
+ ### 2. Select Vision Model
24
+ Open the interactive TUI selector to choose your vision describer model from your connected providers:
47
25
  ```bash
48
- pi install git:github.com/bismawy/pi-vision-watcher
26
+ /vision-watcher
49
27
  ```
28
+ *(You can also set it directly: `/vision-watcher model openai/gpt-4o`)*
50
29
 
51
- Then run `/reload` in Pi (or restart Pi).
30
+ ### 3. Work Seamlessly
31
+ Switch to any text-only model in Pi (e.g. DeepSeek, Claude text-only, local models). Whenever you paste an image, attach a file, or the agent runs `read` on an image, `pi-vision-watcher` describes it automatically in the background.
52
32
 
53
33
  ---
54
34
 
55
- ## 🎮 Usage
35
+ ## 🚀 Key Capabilities
36
+
37
+ - đŸŽ¯ **Connected-Only Interactive Picker:** Shows only vision-capable models from providers where you actually have active credentials (`/login`, `models.json`, or environment variables).
38
+ - ⚡ **DataLoader Batching & SHA-256 Cache:** Automatically groups multiple images across parallel tool calls or multi-file prompts into a single batched describer request. Cached images are never re-described.
39
+ - đŸ›Ąī¸ **Proactive False-Vision Healing:** Aggregator providers often mistakenly flag models (like DeepSeek V4) as multimodal, causing HTTP 400 errors (`This model does not support image`). `pi-vision-watcher` proactively forces handoff for these models and auto-heals `models.json` `modelOverrides` in-process.
40
+ - 🧠 **Thinking & Reasoning Controls:** Adjust reasoning levels (`off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) for reasoning-capable vision models (o-series, Claude, DeepSeek).
41
+ - 🔄 **Multi-Model Fallback Chains:** Automatically falls back to backup vision models if your primary provider is rate-limited or unavailable.
56
42
 
57
- ### Interactive Commands
43
+ ---
44
+
45
+ ## đŸ•šī¸ Command Reference
58
46
 
59
- | Command | Description |
47
+ | Command | Action |
60
48
  |---|---|
61
- | `/vision-watcher` | Open the interactive TUI picker to select your vision model |
62
- | `/vision-watcher model <provider/id>` | Set the vision describer model directly |
63
- | `/vision-watcher status` | View active configuration and handoff status |
64
- | `/vision-watcher enable` / `disable` | Toggle extension on or off |
65
- | `/vision-watcher auto on` / `off` | Toggle automatic handoff for all non-vision models |
66
- | `/vision-watcher add <provider/id>` | Force handoff for a specific model (e.g. weak vision models) |
49
+ | `/vision-watcher` | Open interactive TUI picker for connected vision models |
50
+ | `/vision-watcher model <provider/id>` | Set primary vision describer directly |
51
+ | `/vision-watcher status` | View current configuration and active model status |
52
+ | `/vision-watcher auto <on\|off>` | Toggle automatic handoff for non-vision models (default: `on`) |
53
+ | `/vision-watcher add <provider/id>` | Force handoff on a specific model |
67
54
  | `/vision-watcher remove <provider/id>` | Remove model from forced handoff list |
68
- | `/vision-watcher thinking <level>` | Set describer thinking level (`off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) |
69
- | `/vision-watcher timeout <ms>` | Set the base per-image description timeout in ms (default 45000) |
70
- | `/vision-watcher prewarm on` / `off` | Enable paste-time prewarming in TUI editor |
71
- | `/vision-watcher fallback on` / `off` | Enable async pasted-path description injection |
72
- | `/vision-watcher clear` | Clear configured vision model |
73
- | `/vision-watcher help` | Display full command reference |
55
+ | `/vision-watcher thinking <level>` | Configure reasoning effort for vision models |
56
+ | `/vision-watcher enable` / `disable` | Toggle extension active state |
57
+ | `/vision-watcher help` | Show in-CLI command documentation |
74
58
 
75
59
  ---
76
60
 
77
- ## âš™ī¸ Configuration
61
+ ## 📖 Deep Dive & Advanced Configuration
62
+
63
+ <details>
64
+ <summary><b>âš™ī¸ Configuration File Schema (<code>pi-vision-watcher.json</code>)</b></summary>
78
65
 
79
66
  Configuration is stored at `~/.pi/agent/extensions/pi-vision-watcher.json`:
80
67
 
@@ -96,48 +83,58 @@ Configuration is stored at `~/.pi/agent/extensions/pi-vision-watcher.json`:
96
83
  }
97
84
  ```
98
85
 
99
- | Field | Default | Description |
100
- |---|---|---|
101
- | `enabled` | `true` | Master switch for vision handoff. |
102
- | `visionModel` | `null` | Primary describer as `provider/id` (`null` = handoff inactive). |
103
- | `fallbackModels` | `[]` | List of fallback `provider/id` models tried in order if primary fails. |
104
- | `autoHandoff` | `true` | Automatically apply handoff to all models lacking native vision. |
105
- | `handoffModels` | `[]` | Additional models forced to receive handoff even if vision-capable. |
106
- | `thinking` / `thinkingLevel` | `false` / `"medium"` | Reasoning effort for vision models that support thinking. |
107
- | `describeTimeoutMs` | `45000` | Base per-image timeout in ms before failing over to fallback models. |
108
- | `prewarmPastedImages` | `false` | Describe images immediately upon pasting into the prompt. |
109
- | `asyncClipboardHandoff` | `false` | Asynchronous injection fallback for pasted image paths. |
110
- | `maxTokens` | `null` | Output token cap for descriptions (`null` = model default). |
111
- | `cacheMax` | `50` | Maximum number of described images cached per session. |
112
- | `maxDescriptionLines` | `0` | Truncate description lines (`0` = unbounded). |
113
-
114
- ---
115
-
116
- ## 🔍 Troubleshooting & Logs
86
+ | Field | Type | Default | Description |
87
+ |---|---|---|---|
88
+ | `enabled` | `boolean` | `true` | Master switch for handoff processing. |
89
+ | `visionModel` | `string \| null` | `null` | Primary describer model ref (`provider/id`). |
90
+ | `fallbackModels` | `string[]` | `[]` | Ordered backup models if the primary model fails. |
91
+ | `autoHandoff` | `boolean` | `true` | Automatically describe images for models lacking native vision. |
92
+ | `handoffModels` | `string[]` | `[]` | Specific model IDs forced to receive descriptions. |
93
+ | `thinking` | `boolean` | `false` | Enable reasoning tokens for vision model. |
94
+ | `thinkingLevel` | `string` | `"medium"` | Reasoning effort (`minimal`, `low`, `medium`, `high`, `xhigh`, `max`). |
95
+ | `describeTimeoutMs` | `number` | `45000` | Per-batch timeout before aborting or triggering fallbacks. |
96
+ | `prewarmPastedImages` | `boolean` | `false` | Start describing clipboard images immediately upon pasting in prompt. |
97
+ | `asyncClipboardHandoff` | `boolean` | `false` | Async clipboard injection fallback mechanism. |
98
+ | `maxTokens` | `number \| null` | `null` | Max output tokens for descriptions (`null` = model default). |
99
+ | `cacheMax` | `number` | `50` | Maximum cached image hashes per session. |
100
+ | `maxDescriptionLines` | `number` | `0` | Truncate lines in description block (`0` = full description). |
101
+
102
+ </details>
103
+
104
+ <details>
105
+ <summary><b>🔍 Troubleshooting, Recovery & Diagnostics</b></summary>
106
+
107
+ ### Structured Error Logging
108
+ If a vision call fails, errors are appended with stack traces and request metadata to:
109
+ ```text
110
+ ~/.pi/agent/logs/pi-vision-watcher/errors.log
111
+ ```
112
+ Failures degrade gracefully to `[Image: description unavailable]` without breaking the agent turn.
117
113
 
118
- - **Error Logs:** Detailed failure traces, timestamps, and config snapshots are recorded in `~/.pi/agent/logs/pi-vision-watcher/errors.log`.
119
- - **Transient Retries:** Failed descriptions are never cached permanently — the next turn automatically re-attempts description generation.
114
+ ### False-Vision Auto-Recovery
115
+ When a model falsely advertises image capability and returns an HTTP 400 rejection:
116
+ 1. `pi-vision-watcher` captures the error in the `message_end` event.
117
+ 2. It automatically updates `~/.pi/agent/models.json` under `providers.<name>.modelOverrides.<model>.input = ["text"]`.
118
+ 3. It triggers an in-process registry refresh so subsequent turns use handoff naturally.
120
119
 
121
- ---
120
+ </details>
122
121
 
123
- ## đŸ› ī¸ Development
122
+ <details>
123
+ <summary><b>đŸ› ī¸ Development & Testing</b></summary>
124
124
 
125
125
  ```bash
126
- pnpm install
127
- pnpm test # Run Vitest unit tests
128
- pnpm typecheck # Run TypeScript type check
126
+ bun install
127
+ bun run test # Run Vitest test suite (240+ unit tests)
128
+ bun run typecheck # Run TypeScript compiler check
129
+ bun run lint:dead # Scan for unused exports with Knip
129
130
  ```
130
131
 
131
- ---
132
+ </details>
132
133
 
133
- ## 📜 Credits & License
134
-
135
- `pi-vision-watcher` is inspired by and forked from [`pi-vision-handoff`](https://github.com/monotykamary/pi-vision-handoff) by [Tom X Nguyen](https://github.com/monotykamary) (originating from the concept in `pi-umans-provider`).
134
+ ---
136
135
 
137
- **Key Enhancements:**
138
- - Filters picker to only authenticated/connected models.
139
- - Added thinking & reasoning controls for modern reasoning vision models.
140
- - Multi-model fallback chain support.
141
- - Streamlined settings and UI badging.
136
+ ## 📜 License & Acknowledgments
142
137
 
143
- Released under the [MIT License](./LICENSE).
138
+ - Built for the **[pi coding agent](https://github.com/earendil-works/pi-coding-agent)** ecosystem.
139
+ - Evolved from concepts in `pi-vision-handoff` by Tom X Nguyen and `pi-umans-provider`.
140
+ - Distributed under the **[MIT License](./LICENSE)**.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@bismawy/pi-vision-watcher",
3
- "version": "1.0.8",
3
+ "version": "1.0.9",
4
4
  "description": "Give text-only pi models vision — describe images with a vision model you pick via an interactive picker, then hand off the text description to non-vision models",
5
5
  "type": "module",
6
6
  "author": "bismawy",
@@ -57,7 +57,7 @@
57
57
  "extensions": [
58
58
  "./vision-watcher.ts"
59
59
  ],
60
- "image": "https://raw.githubusercontent.com/bismawy/pi-vision-watcher/main/assets/screenshot.png"
60
+ "image": "https://raw.githubusercontent.com/bismawy/pi-vision-watcher/main/assets/screenshot.webp"
61
61
  },
62
62
  "peerDependencies": {
63
63
  "@earendil-works/pi-ai": "*",