@bismawy/pi-vision-watcher 1.0.9 → 1.0.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +140 -140
  2. package/package.json +1 -1
  3. package/src/index.ts +9 -5
package/README.md CHANGED
@@ -1,140 +1,140 @@
1
- # đŸ‘ī¸ @bismawy/pi-vision-watcher
2
-
3
- **Give text-only [pi](https://github.com/earendil-works/pi-coding-agent) models vision capabilities.**
4
-
5
- Seamlessly inspect, describe, and convert visual inputs (screenshots, mockups, terminal errors, clipboard pastes) into structured descriptions using your preferred vision model, and hand them off to text-only coding models without interrupting your workflow.
6
-
7
- [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
8
- [![npm](https://img.shields.io/npm/v/@bismawy/pi-vision-watcher)](https://www.npmjs.com/package/@bismawy/pi-vision-watcher)
9
- [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
10
-
11
- ![pi-vision-watcher](./assets/screenshot.webp)
12
-
13
- ---
14
-
15
- ## ⚡ Quick Start
16
-
17
- ### 1. Installation
18
- ```bash
19
- pi install npm:@bismawy/pi-vision-watcher
20
- ```
21
- *(Or install directly from Git: `pi install git:github.com/bismawy/pi-vision-watcher`)*
22
-
23
- ### 2. Select Vision Model
24
- Open the interactive TUI selector to choose your vision describer model from your connected providers:
25
- ```bash
26
- /vision-watcher
27
- ```
28
- *(You can also set it directly: `/vision-watcher model openai/gpt-4o`)*
29
-
30
- ### 3. Work Seamlessly
31
- Switch to any text-only model in Pi (e.g. DeepSeek, Claude text-only, local models). Whenever you paste an image, attach a file, or the agent runs `read` on an image, `pi-vision-watcher` describes it automatically in the background.
32
-
33
- ---
34
-
35
- ## 🚀 Key Capabilities
36
-
37
- - đŸŽ¯ **Connected-Only Interactive Picker:** Shows only vision-capable models from providers where you actually have active credentials (`/login`, `models.json`, or environment variables).
38
- - ⚡ **DataLoader Batching & SHA-256 Cache:** Automatically groups multiple images across parallel tool calls or multi-file prompts into a single batched describer request. Cached images are never re-described.
39
- - đŸ›Ąī¸ **Proactive False-Vision Healing:** Aggregator providers often mistakenly flag models (like DeepSeek V4) as multimodal, causing HTTP 400 errors (`This model does not support image`). `pi-vision-watcher` proactively forces handoff for these models and auto-heals `models.json` `modelOverrides` in-process.
40
- - 🧠 **Thinking & Reasoning Controls:** Adjust reasoning levels (`off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) for reasoning-capable vision models (o-series, Claude, DeepSeek).
41
- - 🔄 **Multi-Model Fallback Chains:** Automatically falls back to backup vision models if your primary provider is rate-limited or unavailable.
42
-
43
- ---
44
-
45
- ## đŸ•šī¸ Command Reference
46
-
47
- | Command | Action |
48
- |---|---|
49
- | `/vision-watcher` | Open interactive TUI picker for connected vision models |
50
- | `/vision-watcher model <provider/id>` | Set primary vision describer directly |
51
- | `/vision-watcher status` | View current configuration and active model status |
52
- | `/vision-watcher auto <on\|off>` | Toggle automatic handoff for non-vision models (default: `on`) |
53
- | `/vision-watcher add <provider/id>` | Force handoff on a specific model |
54
- | `/vision-watcher remove <provider/id>` | Remove model from forced handoff list |
55
- | `/vision-watcher thinking <level>` | Configure reasoning effort for vision models |
56
- | `/vision-watcher enable` / `disable` | Toggle extension active state |
57
- | `/vision-watcher help` | Show in-CLI command documentation |
58
-
59
- ---
60
-
61
- ## 📖 Deep Dive & Advanced Configuration
62
-
63
- <details>
64
- <summary><b>âš™ī¸ Configuration File Schema (<code>pi-vision-watcher.json</code>)</b></summary>
65
-
66
- Configuration is stored at `~/.pi/agent/extensions/pi-vision-watcher.json`:
67
-
68
- ```json
69
- {
70
- "enabled": true,
71
- "visionModel": "openai/gpt-4o",
72
- "fallbackModels": [],
73
- "autoHandoff": true,
74
- "handoffModels": [],
75
- "thinking": false,
76
- "thinkingLevel": "medium",
77
- "describeTimeoutMs": 45000,
78
- "prewarmPastedImages": false,
79
- "asyncClipboardHandoff": false,
80
- "maxTokens": null,
81
- "cacheMax": 50,
82
- "maxDescriptionLines": 0
83
- }
84
- ```
85
-
86
- | Field | Type | Default | Description |
87
- |---|---|---|---|
88
- | `enabled` | `boolean` | `true` | Master switch for handoff processing. |
89
- | `visionModel` | `string \| null` | `null` | Primary describer model ref (`provider/id`). |
90
- | `fallbackModels` | `string[]` | `[]` | Ordered backup models if the primary model fails. |
91
- | `autoHandoff` | `boolean` | `true` | Automatically describe images for models lacking native vision. |
92
- | `handoffModels` | `string[]` | `[]` | Specific model IDs forced to receive descriptions. |
93
- | `thinking` | `boolean` | `false` | Enable reasoning tokens for vision model. |
94
- | `thinkingLevel` | `string` | `"medium"` | Reasoning effort (`minimal`, `low`, `medium`, `high`, `xhigh`, `max`). |
95
- | `describeTimeoutMs` | `number` | `45000` | Per-batch timeout before aborting or triggering fallbacks. |
96
- | `prewarmPastedImages` | `boolean` | `false` | Start describing clipboard images immediately upon pasting in prompt. |
97
- | `asyncClipboardHandoff` | `boolean` | `false` | Async clipboard injection fallback mechanism. |
98
- | `maxTokens` | `number \| null` | `null` | Max output tokens for descriptions (`null` = model default). |
99
- | `cacheMax` | `number` | `50` | Maximum cached image hashes per session. |
100
- | `maxDescriptionLines` | `number` | `0` | Truncate lines in description block (`0` = full description). |
101
-
102
- </details>
103
-
104
- <details>
105
- <summary><b>🔍 Troubleshooting, Recovery & Diagnostics</b></summary>
106
-
107
- ### Structured Error Logging
108
- If a vision call fails, errors are appended with stack traces and request metadata to:
109
- ```text
110
- ~/.pi/agent/logs/pi-vision-watcher/errors.log
111
- ```
112
- Failures degrade gracefully to `[Image: description unavailable]` without breaking the agent turn.
113
-
114
- ### False-Vision Auto-Recovery
115
- When a model falsely advertises image capability and returns an HTTP 400 rejection:
116
- 1. `pi-vision-watcher` captures the error in the `message_end` event.
117
- 2. It automatically updates `~/.pi/agent/models.json` under `providers.<name>.modelOverrides.<model>.input = ["text"]`.
118
- 3. It triggers an in-process registry refresh so subsequent turns use handoff naturally.
119
-
120
- </details>
121
-
122
- <details>
123
- <summary><b>đŸ› ī¸ Development & Testing</b></summary>
124
-
125
- ```bash
126
- bun install
127
- bun run test # Run Vitest test suite (240+ unit tests)
128
- bun run typecheck # Run TypeScript compiler check
129
- bun run lint:dead # Scan for unused exports with Knip
130
- ```
131
-
132
- </details>
133
-
134
- ---
135
-
136
- ## 📜 License & Acknowledgments
137
-
138
- - Built for the **[pi coding agent](https://github.com/earendil-works/pi-coding-agent)** ecosystem.
139
- - Evolved from concepts in `pi-vision-handoff` by Tom X Nguyen and `pi-umans-provider`.
140
- - Distributed under the **[MIT License](./LICENSE)**.
1
+ # đŸ‘ī¸ @bismawy/pi-vision-watcher
2
+
3
+ **Give text-only [pi](https://github.com/earendil-works/pi-coding-agent) models vision capabilities.**
4
+
5
+ Seamlessly inspect, describe, and convert visual inputs (screenshots, mockups, terminal errors, clipboard pastes) into structured descriptions using your preferred vision model, and hand them off to text-only coding models without interrupting your workflow.
6
+
7
+ [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
8
+ [![npm](https://img.shields.io/npm/v/@bismawy/pi-vision-watcher)](https://www.npmjs.com/package/@bismawy/pi-vision-watcher)
9
+ [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
10
+
11
+ ![pi-vision-watcher](https://raw.githubusercontent.com/bismawy/pi-vision-watcher/main/assets/screenshot.webp)
12
+
13
+ ---
14
+
15
+ ## ⚡ Quick Start
16
+
17
+ ### 1. Installation
18
+ ```bash
19
+ pi install npm:@bismawy/pi-vision-watcher
20
+ ```
21
+ *(Or install directly from Git: `pi install git:github.com/bismawy/pi-vision-watcher`)*
22
+
23
+ ### 2. Select Vision Model
24
+ Open the interactive TUI selector to choose your vision describer model from your connected providers:
25
+ ```bash
26
+ /vision-watcher
27
+ ```
28
+ *(You can also set it directly: `/vision-watcher model openai/gpt-4o`)*
29
+
30
+ ### 3. Work Seamlessly
31
+ Switch to any text-only model in Pi (e.g. DeepSeek, Claude text-only, local models). Whenever you paste an image, attach a file, or the agent runs `read` on an image, `pi-vision-watcher` describes it automatically in the background.
32
+
33
+ ---
34
+
35
+ ## 🚀 Key Capabilities
36
+
37
+ - đŸŽ¯ **Connected-Only Interactive Picker:** Shows only vision-capable models from providers where you actually have active credentials (`/login`, `models.json`, or environment variables).
38
+ - ⚡ **DataLoader Batching & SHA-256 Cache:** Automatically groups multiple images across parallel tool calls or multi-file prompts into a single batched describer request. Cached images are never re-described.
39
+ - đŸ›Ąī¸ **Proactive False-Vision Healing:** Aggregator providers often mistakenly flag models (like DeepSeek V4) as multimodal, causing HTTP 400 errors (`This model does not support image`). `pi-vision-watcher` proactively forces handoff for these models and auto-heals `models.json` `modelOverrides` in-process.
40
+ - 🧠 **Thinking & Reasoning Controls:** Adjust reasoning levels (`off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max`) for reasoning-capable vision models (o-series, Claude, DeepSeek).
41
+ - 🔄 **Multi-Model Fallback Chains:** Automatically falls back to backup vision models if your primary provider is rate-limited or unavailable.
42
+
43
+ ---
44
+
45
+ ## đŸ•šī¸ Command Reference
46
+
47
+ | Command | Action |
48
+ |---|---|
49
+ | `/vision-watcher` | Open interactive TUI picker for connected vision models |
50
+ | `/vision-watcher model <provider/id>` | Set primary vision describer directly |
51
+ | `/vision-watcher status` | View current configuration and active model status |
52
+ | `/vision-watcher auto <on\|off>` | Toggle automatic handoff for non-vision models (default: `on`) |
53
+ | `/vision-watcher add <provider/id>` | Force handoff on a specific model |
54
+ | `/vision-watcher remove <provider/id>` | Remove model from forced handoff list |
55
+ | `/vision-watcher thinking <level>` | Configure reasoning effort for vision models |
56
+ | `/vision-watcher enable` / `disable` | Toggle extension active state |
57
+ | `/vision-watcher help` | Show in-CLI command documentation |
58
+
59
+ ---
60
+
61
+ ## 📖 Deep Dive & Advanced Configuration
62
+
63
+ <details>
64
+ <summary><b>âš™ī¸ Configuration File Schema (<code>pi-vision-watcher.json</code>)</b></summary>
65
+
66
+ Configuration is stored at `~/.pi/agent/extensions/pi-vision-watcher.json`:
67
+
68
+ ```json
69
+ {
70
+ "enabled": true,
71
+ "visionModel": "openai/gpt-4o",
72
+ "fallbackModels": [],
73
+ "autoHandoff": true,
74
+ "handoffModels": [],
75
+ "thinking": false,
76
+ "thinkingLevel": "medium",
77
+ "describeTimeoutMs": 45000,
78
+ "prewarmPastedImages": false,
79
+ "asyncClipboardHandoff": false,
80
+ "maxTokens": null,
81
+ "cacheMax": 50,
82
+ "maxDescriptionLines": 0
83
+ }
84
+ ```
85
+
86
+ | Field | Type | Default | Description |
87
+ |---|---|---|---|
88
+ | `enabled` | `boolean` | `true` | Master switch for handoff processing. |
89
+ | `visionModel` | `string \| null` | `null` | Primary describer model ref (`provider/id`). |
90
+ | `fallbackModels` | `string[]` | `[]` | Ordered backup models if the primary model fails. |
91
+ | `autoHandoff` | `boolean` | `true` | Automatically describe images for models lacking native vision. |
92
+ | `handoffModels` | `string[]` | `[]` | Specific model IDs forced to receive descriptions. |
93
+ | `thinking` | `boolean` | `false` | Enable reasoning tokens for vision model. |
94
+ | `thinkingLevel` | `string` | `"medium"` | Reasoning effort (`minimal`, `low`, `medium`, `high`, `xhigh`, `max`). |
95
+ | `describeTimeoutMs` | `number` | `45000` | Per-batch timeout before aborting or triggering fallbacks. |
96
+ | `prewarmPastedImages` | `boolean` | `false` | Start describing clipboard images immediately upon pasting in prompt. |
97
+ | `asyncClipboardHandoff` | `boolean` | `false` | Async clipboard injection fallback mechanism. |
98
+ | `maxTokens` | `number \| null` | `null` | Max output tokens for descriptions (`null` = model default). |
99
+ | `cacheMax` | `number` | `50` | Maximum cached image hashes per session. |
100
+ | `maxDescriptionLines` | `number` | `0` | Truncate lines in description block (`0` = full description). |
101
+
102
+ </details>
103
+
104
+ <details>
105
+ <summary><b>🔍 Troubleshooting, Recovery & Diagnostics</b></summary>
106
+
107
+ ### Structured Error Logging
108
+ If a vision call fails, errors are appended with stack traces and request metadata to:
109
+ ```text
110
+ ~/.pi/agent/logs/pi-vision-watcher/errors.log
111
+ ```
112
+ Failures degrade gracefully to `[Image: description unavailable]` without breaking the agent turn.
113
+
114
+ ### False-Vision Auto-Recovery
115
+ When a model falsely advertises image capability and returns an HTTP 400 rejection:
116
+ 1. `pi-vision-watcher` captures the error in the `message_end` event.
117
+ 2. It automatically updates `~/.pi/agent/models.json` under `providers.<name>.modelOverrides.<model>.input = ["text"]`.
118
+ 3. It triggers an in-process registry refresh so subsequent turns use handoff naturally.
119
+
120
+ </details>
121
+
122
+ <details>
123
+ <summary><b>đŸ› ī¸ Development & Testing</b></summary>
124
+
125
+ ```bash
126
+ bun install
127
+ bun run test # Run Vitest test suite (240+ unit tests)
128
+ bun run typecheck # Run TypeScript compiler check
129
+ bun run lint:dead # Scan for unused exports with Knip
130
+ ```
131
+
132
+ </details>
133
+
134
+ ---
135
+
136
+ ## 📜 License & Acknowledgments
137
+
138
+ - Built for the **[pi coding agent](https://github.com/earendil-works/pi-coding-agent)** ecosystem.
139
+ - Evolved from concepts in `pi-vision-handoff` by Tom X Nguyen and `pi-umans-provider`.
140
+ - Distributed under the **[MIT License](./LICENSE)**.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@bismawy/pi-vision-watcher",
3
- "version": "1.0.9",
3
+ "version": "1.0.10",
4
4
  "description": "Give text-only pi models vision — describe images with a vision model you pick via an interactive picker, then hand off the text description to non-vision models",
5
5
  "type": "module",
6
6
  "author": "bismawy",
package/src/index.ts CHANGED
@@ -242,20 +242,24 @@ export function formatModelRef(provider: string, id: string): string {
242
242
  * passes the autoHandoff vision check, so the raw image reaches the provider
243
243
  * and 400s — we learn from that error and force handoff for the model. */
244
244
  export function isImageNotSupportedError(text: string): boolean {
245
- return /not support (?:the )?image|image[s]? (?:input[s]? )?(?:is |are )?not supported|does not accept image|unsupported (?:image|multimodal)/i.test(
245
+ return /not support (?:the )?image|image[s]? (?:input[s]? )?(?:is |are )?not supported|does not accept image|unsupported (?:image|multimodal)|only (?:supports|accepts) text|unsupported content type[\s\S]{0,40}image/i.test(
246
246
  text,
247
247
  );
248
248
  }
249
249
 
250
250
  /** Models whose registry entries commonly declare image input while the
251
- * backend rejects images (DeepSeek V4 via aggregator proxies). Conservative:
252
- * VL / Vision / Janus variants keep image input. Used so autoHandoff covers
253
- * them on the FIRST send — waiting for a 400 is too late. */
251
+ * backend rejects images (DeepSeek V4, GLM 4/5 non-V via aggregator
252
+ * proxies). Conservative: VL / Vision / Janus / *V variants keep image
253
+ * input. Used so autoHandoff covers them on the FIRST send — waiting for a
254
+ * 400 is too late. */
254
255
  export function isKnownTextOnlyFalselyVision(id: string | undefined | null): boolean {
255
256
  if (!id) return false;
256
257
  const n = id.toLowerCase();
257
258
  if (n.includes("vl") || n.includes("vision") || n.includes("janus")) return false;
258
- return /deepseek[-_./]*v4/.test(n);
259
+ if (/\dv(?=[-_./\s]|$)/.test(n)) return false; // glm-4v, glm-4.5v — vision variants
260
+ if (/deepseek[-_./]*v4/.test(n)) return true;
261
+ if (/glm[-_./]*[45]/.test(n)) return true;
262
+ return false;
259
263
  }
260
264
 
261
265
  /** Extract the provider error string from a finalized assistant message.