@anionex/dsh-vision-toolkit 0.1.32 → 0.1.33
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.i18n.yaml +2 -2
- package/README.md +53 -158
- package/README.zh.md +58 -166
- package/assets/upstream/README.md +4 -2
- package/assets/upstream/focus-hint-comparison-1.webp +0 -0
- package/assets/upstream/focus-hint-comparison-2.webp +0 -0
- package/assets/upstream/ui-fast-restore-reference.webp +0 -0
- package/assets/upstream/ui-fast-restore-result.webp +0 -0
- package/docs/python-runtime.i18n.yaml +6 -0
- package/docs/python-runtime.md +75 -0
- package/docs/python-runtime.zh.md +75 -0
- package/lib/client.js +270 -4
- package/lib/client.js.map +1 -1
- package/lib/config.js +2 -0
- package/lib/config.js.map +1 -1
- package/lib/image-input-variants.js +18 -5
- package/lib/image-input-variants.js.map +1 -1
- package/lib/index.js +1 -1
- package/lib/index.js.map +1 -1
- package/lib/runtime-manager.js +10 -3
- package/lib/runtime-manager.js.map +1 -1
- package/lib/types/client/display-config.d.ts +24 -0
- package/lib/types/client/display-config.d.ts.map +1 -0
- package/lib/types/client/index.d.ts +10 -0
- package/lib/types/client/index.d.ts.map +1 -1
- package/lib/types/client/model-variants-hider.d.ts +40 -0
- package/lib/types/client/model-variants-hider.d.ts.map +1 -0
- package/lib/types/client/paste-images.d.ts.map +1 -1
- package/lib/types/config.d.ts +10 -0
- package/lib/types/config.d.ts.map +1 -1
- package/lib/types/image-input-variants.d.ts +2 -1
- package/lib/types/image-input-variants.d.ts.map +1 -1
- package/lib/types/index.d.ts.map +1 -1
- package/lib/types/runtime-manager.d.ts.map +1 -1
- package/lib/types/web.d.ts +17 -1
- package/lib/types/web.d.ts.map +1 -1
- package/lib/web.js +36 -1
- package/lib/web.js.map +1 -1
- package/package.json +1 -1
- package/src/client/display-config.ts +62 -0
- package/src/client/index.tsx +44 -2
- package/src/client/model-variants-hider.ts +159 -0
- package/src/client/paste-images.tsx +5 -1
- package/src/config.ts +12 -0
- package/src/image-input-variants.ts +16 -3
- package/src/index.ts +1 -0
- package/src/runtime-manager.ts +10 -2
- package/src/web.ts +39 -0
- package/assets/upstream/image-qa.webp +0 -0
- package/assets/upstream/screenshot-debugging.webp +0 -0
package/README.i18n.yaml
CHANGED
|
@@ -2,5 +2,5 @@
|
|
|
2
2
|
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
|
3
3
|
# after editing either side, bring the other along and re-record with:
|
|
4
4
|
# pnpm run verify-translation-pairing --write dsh-vision-toolkit/README.md
|
|
5
|
-
README.md:
|
|
6
|
-
README.zh.md:
|
|
5
|
+
README.md: 5083c3bc9c77ab79eed171c5fbdeb428df274f77
|
|
6
|
+
README.zh.md: 979935b0fad2fcdeccf290372dcb5b12176a2ecf
|
package/README.md
CHANGED
|
@@ -16,33 +16,27 @@
|
|
|
16
16
|
[](LICENSE)
|
|
17
17
|
[](cordis.patch.yml)
|
|
18
18
|
|
|
19
|
-
|
|
20
19
|
**A more powerful vision toolkit—give text-only models in DeepSeek Harness eyes: image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks in one toolkit and Skill.**
|
|
21
20
|
|
|
22
21
|
🚀 Paste an image and ask directly | Install with one command | Built-in free vision | Broad use cases
|
|
23
22
|
|
|
24
|
-
|
|
25
|
-
<a href="#highlights">Highlights</a> | <a href="#quick-start-three-steps">Quick start</a> | <a href="#common-workflows">Common workflows</a> | <a href="#toolbox">Toolbox</a> | <a href="#configuration-and-limits">Configuration</a> | <a href="#troubleshooting">Troubleshooting</a> | <a href="#development-and-community">Community</a>
|
|
26
|
-
</p>
|
|
23
|
+
[Highlights](#highlights) | [Quick start](#quick-start-three-steps) | [Toolbox](#toolbox) | [Configuration and limits](#configuration-and-limits) | [Troubleshooting](#troubleshooting) | [Community](#development-and-community)
|
|
27
24
|
|
|
28
25
|
🌐 **English** | [中文](README.zh.md)
|
|
29
26
|
|
|
30
27
|
</div>
|
|
31
28
|
|
|
32
|
-
If you use DeepSeek or another text-only model in DeepSeek Harness (DSH), you may have run into the same problems: the model cannot see a screenshot, generic descriptions miss the point, buttons have no usable coordinates, and a rebuilt page can look “close enough” without a way to measure the remaining difference.
|
|
33
|
-
|
|
34
29
|
🏆 This project is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem: it was initiated before internal beta and built during the beta with reference to [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit).
|
|
35
30
|
|
|
36
31
|
> **Original work:** The system and division of responsibilities behind these visual tools, together with the `vision-skills` Skill, were personally created and continuously refined by the author through long-term real-world use and repeated iteration.
|
|
37
32
|
|
|
38
33
|
## Highlights
|
|
39
34
|
|
|
40
|
-
- **Paste and
|
|
41
|
-
- **A seamless image workflow.** Native thumbnails, session history, and workspace paths stay intact; Web can preview artifacts and Headless can continue using the same structured results.
|
|
35
|
+
- **Paste an image and ask directly.** In DSH Web, pasting an image switches the text-only model to its `(Vision Toolkit)` variant automatically — no manual path copying or model changes. Native thumbnails, session history, and workspace paths stay intact; Web can preview artifacts.
|
|
42
36
|
- **One command to install.** The built-in free Gemini 3.7 Flash vision service is ready after installation, with no API key required.
|
|
43
|
-
- **Built-in free vision.** The shared service works immediately after installation with a quota of **
|
|
44
|
-
- **
|
|
45
|
-
- **A
|
|
37
|
+
- **Built-in free vision quota.** The shared service works immediately after installation with a quota of **100 images per machine per day**.
|
|
38
|
+
- **Not just a caption — the content that matters.** The model does not produce a generic description; it extracts evidence around the current task, such as “Where is the error?” or “Where is the button?”.
|
|
39
|
+
- **A battle-tested visual-task methodology.** The bundled Skill tells the agent what to look at for different visual tasks, which tool to choose, how to proceed, and how to verify the result.
|
|
46
40
|
|
|
47
41
|
[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) gives an agent more than image captions: it can read, locate, crop, trace, rebuild, and verify visual work. DSH Vision Toolkit is its native DeepSeek Harness integration, bringing that workflow into Web and Headless Profiles.
|
|
48
42
|
|
|
@@ -51,7 +45,7 @@ This project has two layers:
|
|
|
51
45
|
1. **Visual tools and a Skill:** the agent learns when to inspect, ground, OCR, crop, trace, or compare pixels.
|
|
52
46
|
2. **Native DSH integration:** those capabilities live inside Profiles, sessions, Settings, Artifacts, and the Web UI, with a free Gemini 3.7 Flash vision service ready after installation.
|
|
53
47
|
|
|
54
|
-
> **Install and use it immediately.** The default setup includes a free Gemini 3.7 Flash vision service and requires no API key.
|
|
48
|
+
> **Install and use it immediately.** The default setup includes a free Gemini 3.7 Flash vision service and requires no API key.
|
|
55
49
|
|
|
56
50
|
```sh
|
|
57
51
|
dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
@@ -59,22 +53,18 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
|
59
53
|
|
|
60
54
|
**Upstream toolkit:** [Anionex/agent-vision-toolkit](https://github.com/Anionex/agent-vision-toolkit) · **Project website:** [agent-vision.anionex.me](https://agent-vision.anionex.me)
|
|
61
55
|
|
|
62
|
-
|
|
63
|
-
<summary><strong>Table of contents</strong></summary>
|
|
56
|
+
**Contents**
|
|
64
57
|
|
|
65
58
|
- [Highlights](#highlights)
|
|
66
59
|
- [Recent updates](#recent-updates)
|
|
67
60
|
- [Who it is for](#who-it-is-for)
|
|
68
61
|
- [See it in action](#see-it-in-action)
|
|
69
62
|
- [Quick start: three steps](#quick-start-three-steps)
|
|
70
|
-
- [Common workflows](#common-workflows)
|
|
71
63
|
- [Toolbox](#toolbox)
|
|
72
64
|
- [Configuration and limits](#configuration-and-limits)
|
|
73
65
|
- [Troubleshooting](#troubleshooting)
|
|
74
66
|
- [Development and community](#development-and-community)
|
|
75
67
|
|
|
76
|
-
</details>
|
|
77
|
-
|
|
78
68
|
## Recent updates
|
|
79
69
|
|
|
80
70
|
- **2026-08-16 · Windows Python:** Added Microsoft Store Python support, fixing first-time isolated-runtime setup failures for affected Windows users.
|
|
@@ -86,14 +76,18 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
|
86
76
|
|
|
87
77
|
## Who it is for
|
|
88
78
|
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
|
95
|
-
|
|
|
96
|
-
|
|
|
79
|
+
1. Want an interaction experience similar to a multimodal model: paste an image directly and ask a question or make a request.
|
|
80
|
+
2. Want more than image Q&A — complete more complex, high-value visual tasks such as turning a sketch into a front-end page, converting an image into HTML, or extracting chat messages from long screenshots; more scenarios are added over time.
|
|
81
|
+
|
|
82
|
+
The bundled `vision-skills` Skill carries the complete upstream playbooks, explaining when to use each workflow, in what order to call the tools, and how to verify the result:
|
|
83
|
+
|
|
84
|
+
| Playbook | What the agent learns to do |
|
|
85
|
+
| --- | --- |
|
|
86
|
+
| [Read long screenshots, chat histories, and scrolling pages](assets/skill/references/long-screenshot-ocr.md) | Find low-content cut bands, OCR each chunk in order, preserve chat speakers/timestamps/quotes, merge only duplicated overlap, and surface risky boundaries for verification |
|
|
87
|
+
| [Rebuild a UI from a screenshot or design](assets/skill/references/restore-ui.md) | Reuse project components and assets first, then combine code-native UI, extracted visuals, rendered screenshots, and visual comparison to align a page or component |
|
|
88
|
+
| [Restore an icon, logo, illustration, or other graphic](assets/skill/references/restore-graphic.md) | Extract a transparent PNG from the source image, or rebuild an editable/scalable SVG when needed, then verify shape, color, and alpha edges |
|
|
89
|
+
| [Turn a sketch, diagram, or whiteboard into structured code](assets/skill/references/restore-structure.md) | Recover nodes, labels, connections, and directions as editable Mermaid, Graphviz, or another structured representation |
|
|
90
|
+
| [Operate a GUI from screenshots](assets/skill/references/gui.md) | Locate a control, perform one action, capture the screen again, and verify the resulting state before continuing |
|
|
97
91
|
|
|
98
92
|
## See it in action
|
|
99
93
|
|
|
@@ -112,6 +106,8 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
|
112
106
|
<img src="assets/upstream/infographic-result.webp" width="49%" alt="Editable HTML and CSS reconstruction created from the reference screenshot" />
|
|
113
107
|
</p>
|
|
114
108
|
|
|
109
|
+
> Prompt example: “(Use vision-skills) Rebuild this image into HTML.”
|
|
110
|
+
|
|
115
111
|
*Left: the reference screenshot. Right: an editable HTML/CSS result. The result can continue into screenshot rendering and pixel comparison instead of ending as an image description.*
|
|
116
112
|
|
|
117
113
|
### Sketch to working interface
|
|
@@ -123,15 +119,19 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
|
123
119
|
|
|
124
120
|
*Left: a hand-drawn reference. Right: the working interface reconstructed from it.*
|
|
125
121
|
|
|
126
|
-
|
|
122
|
+
> Prompt example: “(Use vision-skills) Turn this sketch into a working front-end page.”
|
|
127
123
|
|
|
128
|
-
|
|
124
|
+
### Fast UI restoration: an approximate first pass
|
|
129
125
|
|
|
130
|
-
<p>
|
|
131
|
-
<img src="
|
|
132
|
-
<img src="
|
|
126
|
+
<p align="center">
|
|
127
|
+
<img src="assets/upstream/ui-fast-restore-reference.webp" width="49%" alt="Original YouMind homepage used as the fast UI restoration reference" />
|
|
128
|
+
<img src="assets/upstream/ui-fast-restore-result.webp" width="49%" alt="Approximate YouMind homepage produced with fast UI restoration mode" />
|
|
133
129
|
</p>
|
|
134
130
|
|
|
131
|
+
> Prompt example: “(Use vision-skills) Quickly rebuild this image into HTML.”
|
|
132
|
+
|
|
133
|
+
*Left: the original page. Right: a fast reconstruction that preserves the main layout, content, and visual hierarchy while allowing approximate colors and library icons. Fast mode targets a first screenshot in about three minutes.*
|
|
134
|
+
|
|
135
135
|
## Quick start: three steps
|
|
136
136
|
|
|
137
137
|
### 1. Install
|
|
@@ -163,23 +163,12 @@ Crop this icon and convert it to SVG.
|
|
|
163
163
|
Rebuild the page from reference.png. After each pass, render it and run a pixel diff until the major differences are gone.
|
|
164
164
|
```
|
|
165
165
|
|
|
166
|
-
## Common workflows
|
|
167
|
-
|
|
168
|
-
| Task | Recommended workflow |
|
|
169
|
-
|---|---|
|
|
170
|
-
| Image Q&A or screenshot debugging | Inspect → answer around the current question → locate details when needed |
|
|
171
|
-
| Find a button, icon, or text region | Ground the target → return pixel box → create a labeled preview |
|
|
172
|
-
| Extract an icon from a screenshot | Ground → crop → trace to SVG |
|
|
173
|
-
| Read a long webpage screenshot | Split → OCR → merge Markdown → audit boundaries |
|
|
174
|
-
| Recreate a page or component | Reference → implementation → HTML screenshot → pixel diff → iterate |
|
|
175
|
-
| Extract brand visuals | Crop region → analyze dominant colors → extract foreground → export transparent PNG |
|
|
176
|
-
|
|
177
166
|
## Toolbox
|
|
178
167
|
|
|
179
168
|
The plugin provides 10 tools that can be called independently or composed into a workflow:
|
|
180
169
|
|
|
181
170
|
| Tool | Best question to ask | Main result |
|
|
182
|
-
|
|
171
|
+
| --- | --- | --- |
|
|
183
172
|
| `vision_glance` | “What is happening in this image?” | Focused answer, description, OCR, or multi-image comparison |
|
|
184
173
|
| `vision_ground` | “Where is the thing I need?” | Original pixel coordinates and optional boxed preview |
|
|
185
174
|
| `vision_detect` | “Which buttons, icons, or elements are present?” | Numbered element inventory, coordinates, and optional preview |
|
|
@@ -197,10 +186,18 @@ For a long HTML document, pass `fullPage=true`. The requested width and height r
|
|
|
197
186
|
|
|
198
187
|
## How it works
|
|
199
188
|
|
|
200
|
-
The plugin keeps image understanding and deterministic local image processing in one Agent workflow.
|
|
189
|
+
The plugin keeps image understanding and deterministic local image processing in one Agent workflow. The diagram below shows the implementation boundary.
|
|
190
|
+
|
|
191
|
+
### Descriptions that keep the task in view
|
|
192
|
+
|
|
193
|
+
Most vision bridges for text-only models ask a multimodal model for a generic description and hand it to the text model, adding a semantic layer where information is lost. Vision Toolkit instead recovers **why the agent wants to look at the image**: the user message or the model's stated reason becomes a focus hint passed to the vision model. The result is a task-aware description that emphasizes what matters for the current step — with fewer tokens, higher accuracy, and faster responses.
|
|
194
|
+
|
|
195
|
+
<p align="center">
|
|
196
|
+
<img src="assets/upstream/focus-hint-comparison-1.webp" width="49%" alt="Generic image descriptions compared with task-aware vision using a focus hint - part 1" />
|
|
197
|
+
<img src="assets/upstream/focus-hint-comparison-2.webp" width="49%" alt="Generic image descriptions compared with task-aware vision using a focus hint - part 2" />
|
|
198
|
+
</p>
|
|
201
199
|
|
|
202
|
-
|
|
203
|
-
<summary><strong>Architecture and image-input behavior</strong></summary>
|
|
200
|
+
**Architecture and image-input behavior**
|
|
204
201
|
|
|
205
202
|
```mermaid
|
|
206
203
|
flowchart LR
|
|
@@ -226,8 +223,6 @@ recorded in `assets/skill/UPSTREAM.json` and `patches/vision-tools-dsh.patch`.
|
|
|
226
223
|
|
|
227
224
|
For routes that DSH positively identifies as text-only, the plugin registers a sibling `<model> (Vision Toolkit)` variant. By default, pasting an image in DSH Web switches to that variant and gives the model both a reusable workspace path and a visual description focused on the current task.
|
|
228
225
|
|
|
229
|
-
</details>
|
|
230
|
-
|
|
231
226
|
## Configuration and limits
|
|
232
227
|
|
|
233
228
|
### Built-in free service
|
|
@@ -245,8 +240,8 @@ Requests that still use the previous `qwen/qwen3.6-27b` model name remain compat
|
|
|
245
240
|
This is a shared zero-configuration entry point, not an unlimited private endpoint. Request safeguards include:
|
|
246
241
|
|
|
247
242
|
| Limit | Current value |
|
|
248
|
-
|
|
249
|
-
| Daily quota |
|
|
243
|
+
| --- | --- |
|
|
244
|
+
| Daily quota | 100 images per machine per day |
|
|
250
245
|
| Images per request | Up to 5 |
|
|
251
246
|
| Image size | 4 MiB per image |
|
|
252
247
|
| Decoded pixels | 20,000,000 per image |
|
|
@@ -278,119 +273,17 @@ OpenAI Chat Completions-compatible endpoints and Anthropic Messages are supporte
|
|
|
278
273
|
|
|
279
274
|
For a trusted internal endpoint that uses a self-signed certificate or MITM proxy, start the DSH process with `VISION_SSL_VERIFY=0`. The plugin forwards that value to the isolated Python runtime; certificate verification remains enabled when the variable is unset or has any other value. The false values `false`, `off`, `no`, `none`, and `disabled` are also accepted, case-insensitively.
|
|
280
275
|
|
|
281
|
-
### Requirements
|
|
282
|
-
|
|
283
|
-
- A DeepSeek Harness Web or Headless Profile.
|
|
284
|
-
- Node.js `^22.19.0` or `>=24.0.0`.
|
|
285
|
-
- Python 3.11+ is usually not needed in advance: the plugin prefers a system Python and otherwise downloads a pinned standalone Python 3.13 automatically, preparing its own isolated environment. Only that first automatic download needs network access.
|
|
286
|
-
- Only `vision_html_screenshot` requires Chrome, Chromium, or Edge.
|
|
287
|
-
- Inputs must be PNG, JPEG, GIF, or WebP files in the session workspace, the platform temporary directory, or an explicitly allowed directory.
|
|
288
|
-
|
|
289
276
|
### Configure the Python runtime
|
|
290
277
|
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
The packaged `managed` runtime creates its own isolated virtual environment. `runtime.python` selects the Python executable used to bootstrap or refresh that environment; it does not replace the managed environment with the interpreter's global site-packages. Set it when automatic discovery fails or when several Python installations exist. The override is also used by `runtime.mode: external`.
|
|
294
|
-
|
|
295
|
-
Python 3.11 or newer is required; the automatically downloaded standalone Python is 3.13.15 and, like a system interpreter, is only used to bootstrap the isolated environment. Without an override, the plugin tries `python3` then `python` on macOS/Linux, and `python`, `py -3`, then `python3` on Windows, before falling back to the standalone download. A configured value is passed as one executable name or path, not as a shell command with arguments, so use `py` (not `py -3`) for the Windows launcher; use an absolute path when you need a specific version.
|
|
278
|
+
Most users never need to configure the Python runtime: the plugin prefers a system Python 3.11+ and otherwise downloads a pinned standalone Python automatically.
|
|
296
279
|
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
```yaml
|
|
300
|
-
- id: vision-toolkit
|
|
301
|
-
config:
|
|
302
|
-
runtime:
|
|
303
|
-
# macOS/Linux system Python
|
|
304
|
-
python: python3
|
|
305
|
-
# Or a project-local environment:
|
|
306
|
-
# python: /absolute/path/to/project/.venv/bin/python
|
|
307
|
-
# Windows venv (forward slashes also work in YAML):
|
|
308
|
-
# python: C:/Users/you/project/.venv/Scripts/python.exe
|
|
309
|
-
# Windows launcher, when its default Python is 3.11+:
|
|
310
|
-
# python: py
|
|
311
|
-
```
|
|
312
|
-
|
|
313
|
-
For a managed runtime, create the project-local interpreter and point `runtime.python` at it. The plugin installs the locked dependencies into its own managed cache, so installing the lockfile into this bootstrap environment is optional:
|
|
314
|
-
|
|
315
|
-
```sh
|
|
316
|
-
python3 --version # must report 3.11 or newer
|
|
317
|
-
uv venv .venv --python 3.13
|
|
318
|
-
```
|
|
319
|
-
|
|
320
|
-
For `runtime.mode: external`, install the locked dependencies using the `runtime/requirements.lock` from the **DSH Vision Toolkit plugin** checkout, then point `runtime.agentVisionToolkitPath` at a separate exact `agent-vision-toolkit` snapshot. The packaged `vendor/agent-vision-toolkit` directory is such a snapshot when it has not been modified:
|
|
321
|
-
|
|
322
|
-
```sh
|
|
323
|
-
uv pip install --python .venv/bin/python \
|
|
324
|
-
-r /absolute/path/to/dsh-vision-toolkit/runtime/requirements.lock
|
|
325
|
-
```
|
|
326
|
-
|
|
327
|
-
```yaml
|
|
328
|
-
- id: vision-toolkit
|
|
329
|
-
config:
|
|
330
|
-
runtime:
|
|
331
|
-
mode: external
|
|
332
|
-
python: /absolute/path/to/dsh-vision-toolkit/.venv/bin/python
|
|
333
|
-
agentVisionToolkitPath: /absolute/path/to/dsh-vision-toolkit/vendor/agent-vision-toolkit
|
|
334
|
-
```
|
|
335
|
-
|
|
336
|
-
On Windows, use `py -3 --version` for the version check and `.venv\Scripts\python.exe` plus `runtime\requirements.lock` in the corresponding commands:
|
|
337
|
-
|
|
338
|
-
```powershell
|
|
339
|
-
py -3 --version # must report 3.11 or newer
|
|
340
|
-
uv venv .venv --python 3.13
|
|
341
|
-
# External mode only; use the plugin checkout's absolute lockfile path:
|
|
342
|
-
uv pip install --python .venv\Scripts\python.exe -r C:\absolute\path\to\dsh-vision-toolkit\runtime\requirements.lock
|
|
343
|
-
```
|
|
344
|
-
|
|
345
|
-
Point `runtime.python` at the same interpreter, save the Profile patch, and restart the Web Profile. Then open **Settings → Vision Toolkit**: the Runtime panel should show the resolved interpreter and Python version, and **Run health check** plus **Test vision model** should complete without the Python-version error. A final smoke test is to place a PNG/JPEG in the session workspace and call `vision_glance`.
|
|
346
|
-
|
|
347
|
-
The path fence automatically allows the session workspace and the platform temporary directory. On macOS/Linux the temporary root is `/tmp`. On Windows it is `TEMP`, then `TMP`, with the operating-system fallback if neither is set; model-generated `/tmp/...` paths are translated to that Windows directory before the normal realpath fence runs. No `allowedDirs` entry is needed for these platform temporary paths.
|
|
348
|
-
|
|
349
|
-
Use `allowedDirs` only for additional trusted input roots outside the workspace and platform temporary directory:
|
|
350
|
-
|
|
351
|
-
```yaml
|
|
352
|
-
- id: vision-toolkit
|
|
353
|
-
config:
|
|
354
|
-
allowedDirs:
|
|
355
|
-
# macOS/Linux example
|
|
356
|
-
- /srv/vision-inputs
|
|
357
|
-
# Windows example (use this instead on Windows)
|
|
358
|
-
# - D:/vision-inputs
|
|
359
|
-
```
|
|
360
|
-
|
|
361
|
-
`allowedDirs` is an input allowlist, not the managed runtime cache. The managed runtime keeps its own files under `$DSH_HOME/cache/dsh-vision-toolkit` (or `~/.dsh/cache/dsh-vision-toolkit` when `DSH_HOME` is unset); that directory does not need to be added. Environment-variable forms such as `$env:TEMP` and `%TEMP%` are not expanded inside `allowedDirs`, so configure extra roots with real absolute paths.
|
|
362
|
-
|
|
363
|
-
<details>
|
|
364
|
-
<summary><strong>Install, upgrade, disable, and uninstall</strong></summary>
|
|
365
|
-
|
|
366
|
-
```sh
|
|
367
|
-
dsh plugin --profile web update @anionex/dsh-vision-toolkit
|
|
368
|
-
dsh plugin --profile web remove @anionex/dsh-vision-toolkit
|
|
369
|
-
```
|
|
370
|
-
|
|
371
|
-
If you are migrating from the retired `@dsh-external/dsh-vision-toolkit`, remove the old package first and install `@anionex/dsh-vision-toolkit`.
|
|
372
|
-
|
|
373
|
-
To disable the bundle temporarily, set this in the Profile patch:
|
|
374
|
-
|
|
375
|
-
```yaml
|
|
376
|
-
- id: vision-toolkit
|
|
377
|
-
disabled: true
|
|
378
|
-
```
|
|
379
|
-
|
|
380
|
-
Restart the Web Profile and refresh the page after enabling or upgrading the Web plugin.
|
|
381
|
-
|
|
382
|
-
</details>
|
|
383
|
-
|
|
384
|
-
### Plugin updates
|
|
385
|
-
|
|
386
|
-
In **Settings → Vision Toolkit**, **Check for updates** queries the Profile's npm registry. For a direct registry installation, **Update and restart** installs only the exact version you confirmed, verifies it, and restarts an explicitly opted-in POSIX Web process on a fixed `--port`. Local/workspace/file/git/URL installs, Windows, dynamic ports, read-only Profiles, and manager-owned processes remain check-only.
|
|
387
|
-
|
|
388
|
-
The updater revalidates the Profile before mutation, snapshots the original manifest and lockfile, and holds a token-owned cross-process lock. The current Web process exits only after the restart helper confirms that the backup is readable and the lock handoff succeeded. When the Profile was already operational, the replacement must report both the target plugin version and a ready runtime; failed replacements restore the original manifest/lockfile and rebuild dependencies with a frozen lockfile before retrying the previous exact version. If automatic recovery itself fails, the backup and lock are preserved and their paths are written to `$DSH_HOME/logs/vision-toolkit-restart.log`. Detached restart requires `DSH_VISION_TOOLKIT_ALLOW_DETACHED_RESTART=1`; unsaved Settings or API-key input blocks installation.
|
|
280
|
+
For advanced setups — overriding `runtime.python`, using `runtime.mode: external`, verifying the runtime, or allowing additional input directories — see [Python runtime configuration](docs/python-runtime.md).
|
|
389
281
|
|
|
390
282
|
## Troubleshooting
|
|
391
283
|
|
|
392
284
|
| Problem | What to do |
|
|
393
|
-
|
|
285
|
+
| --- | --- |
|
|
286
|
+
| The vision-model test fails with `Vision API returned an incompatible response structure` | The base URL usually needs a path prefix. Local OpenAI-compatible services such as LM Studio and Ollama should be entered as `http://127.0.0.1:1234/v1` (include `/v1`); the plugin appends `/chat/completions`, and a port-only address hits an unknown endpoint and returns this error |
|
|
394
287
|
| Pasting an image still says the model does not support image input | Restart the Web Profile, refresh the page, and confirm the selected route has the `(Vision Toolkit)` suffix. You can also place the image in the session workspace and invoke `/vision-skills` |
|
|
395
288
|
| The free service returns 429 | Wait for the `Retry-After` interval, or switch to your own endpoint when you need stable higher volume |
|
|
396
289
|
| The image exceeds a size or pixel limit | Crop or resize it first; the error identifies whether bytes or decoded pixels caused the rejection |
|
|
@@ -399,9 +292,11 @@ The updater revalidates the Profile before mutation, snapshots the original mani
|
|
|
399
292
|
| Chrome is not found | Install Chrome, Chromium, or Edge. Only HTML screenshot rendering is unavailable; the other tools still work |
|
|
400
293
|
| An artifact cannot be previewed | Use **Open file** or the workspace path in the result. Preview URLs exist only while the Web route is available |
|
|
401
294
|
|
|
402
|
-
##
|
|
295
|
+
## FAQ
|
|
296
|
+
|
|
297
|
+
**Will adding a vision model significantly increase costs?**
|
|
403
298
|
|
|
404
|
-
|
|
299
|
+
No. Each inspection sends only the necessary intent and the image to the multimodal model, and context does not accumulate across calls, so the added cost stays small. To reduce it further, a locally deployed small multimodal side model (for example the Gemma 4 or Qwen 3.5/3.6 series) can provide the vision capability.
|
|
405
300
|
|
|
406
301
|
## Development and community
|
|
407
302
|
|
|
@@ -415,7 +310,7 @@ The current release focuses on screenshot understanding, visual grounding, OCR,
|
|
|
415
310
|
<img src="assets/community-group-qr.png" alt="QR code for the agent-vision-toolkit community group" width="240" />
|
|
416
311
|
</p>
|
|
417
312
|
|
|
418
|
-
I'm
|
|
313
|
+
I'm [anionex](https://anionex.me/), an AI-native developer who once ranked **No. 3** on GitHub's global developer trending list, with more than 16k stars across my projects. If you would like to follow my future work, [follow me on GitHub](https://github.com/Anionex).
|
|
419
314
|
|
|
420
315
|
[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) was created by [Anionex](https://anionex.me/). This repository maintains its native DeepSeek Harness integration.
|
|
421
316
|
|