@anionex/dsh-vision-toolkit 0.1.12 → 0.1.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.i18n.yaml +3 -3
- package/README.md +207 -346
- package/README.zh.md +212 -360
- package/docs/requirements-traceability/README.i18n.yaml +2 -2
- package/docs/requirements-traceability/README.md +1 -1
- package/docs/requirements-traceability/README.zh.md +1 -1
- package/lib/client.js +173 -7
- package/lib/client.js.map +1 -1
- package/lib/config.js +4 -0
- package/lib/config.js.map +1 -1
- package/lib/defaults.js +1 -1
- package/lib/defaults.js.map +1 -1
- package/lib/plugin-update.js +1011 -0
- package/lib/plugin-update.js.map +1 -0
- package/lib/runtime-install.js +61 -1
- package/lib/runtime-install.js.map +1 -1
- package/lib/runtime.js +20 -3
- package/lib/runtime.js.map +1 -1
- package/lib/skill.js +7 -3
- package/lib/skill.js.map +1 -1
- package/lib/tools.js +3 -2
- package/lib/tools.js.map +1 -1
- package/lib/types/client/index.d.ts +55 -1
- package/lib/types/client/index.d.ts.map +1 -1
- package/lib/types/config.d.ts.map +1 -1
- package/lib/types/defaults.d.ts +1 -1
- package/lib/types/defaults.d.ts.map +1 -1
- package/lib/types/plugin-update.d.ts +111 -0
- package/lib/types/plugin-update.d.ts.map +1 -0
- package/lib/types/runtime-install.d.ts +11 -0
- package/lib/types/runtime-install.d.ts.map +1 -1
- package/lib/types/runtime.d.ts +3 -0
- package/lib/types/runtime.d.ts.map +1 -1
- package/lib/types/skill.d.ts +1 -1
- package/lib/types/skill.d.ts.map +1 -1
- package/lib/types/tools.d.ts.map +1 -1
- package/lib/types/upstream.d.ts +1 -0
- package/lib/types/upstream.d.ts.map +1 -1
- package/lib/types/web.d.ts +13 -1
- package/lib/types/web.d.ts.map +1 -1
- package/lib/upstream.js +20 -8
- package/lib/upstream.js.map +1 -1
- package/lib/web.js +46 -7
- package/lib/web.js.map +1 -1
- package/package.json +12 -14
- package/src/client/index.tsx +229 -6
- package/src/config.ts +4 -0
- package/src/defaults.ts +1 -1
- package/src/plugin-update.ts +1142 -0
- package/src/runtime-install.ts +77 -1
- package/src/runtime.ts +23 -3
- package/src/skill.ts +7 -3
- package/src/tools.ts +4 -2
- package/src/upstream.ts +21 -8
- package/src/web.ts +69 -3
- package/vendor/agent-vision-toolkit/UPSTREAM_MANIFEST.json +3 -3
- package/vendor/agent-vision-toolkit/bin/crop +0 -0
- package/vendor/agent-vision-toolkit/bin/detect +0 -0
- package/vendor/agent-vision-toolkit/bin/glance +0 -0
- package/vendor/agent-vision-toolkit/bin/ground +0 -0
- package/vendor/agent-vision-toolkit/bin/trace +0 -0
- package/vendor/agent-vision-toolkit/skills/vision-tools/scripts/html_shot.py +323 -11
- package/vendor/agent-vision-toolkit/skills/vision-tools/scripts/long_screenshot_ocr.py +0 -0
- /package/assets/{hero.png → hero-v2.png} +0 -0
package/README.md
CHANGED
|
@@ -1,470 +1,331 @@
|
|
|
1
1
|
<p align="center">
|
|
2
|
-
<img src="assets/hero.png" alt="DSH Vision Toolkit
|
|
2
|
+
<img src="assets/hero-v2.png" alt="DSH Vision Toolkit helps text-only DeepSeek Harness agents understand images and complete visual tasks" />
|
|
3
3
|
</p>
|
|
4
4
|
|
|
5
|
-
<
|
|
5
|
+
<div align="center">
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
English | <a href="https://github.com/Anionex/dsh-vision-toolkit/blob/main/README.zh.md">中文</a>
|
|
9
|
-
</p>
|
|
7
|
+
# DSH Vision Toolkit
|
|
10
8
|
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
<a href="tests"><img src="https://img.shields.io/badge/verified-233%20tests-2EA44F?style=flat-square" alt="Verified: 233 tests" /></a>
|
|
17
|
-
</p>
|
|
9
|
+
[](https://dshfind.com/en/plugins/Anionex/dsh-vision-toolkit)
|
|
10
|
+
[](https://dshfind.com/en/plugins/Anionex/dsh-vision-toolkit)
|
|
11
|
+
[](https://www.npmjs.com/package/@anionex/dsh-vision-toolkit)
|
|
12
|
+
[](LICENSE)
|
|
13
|
+
[](cordis.patch.yml)
|
|
18
14
|
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
15
|
+
**A more powerful vision toolkit—give text-only models in DeepSeek Harness eyes: image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks in one toolkit and Skill.**
|
|
16
|
+
|
|
17
|
+
🚀 Paste an image and ask directly | Install with one command | Built-in free quota | Broad use cases
|
|
18
|
+
|
|
19
|
+
🌐 **English** | [中文](README.zh.md)
|
|
20
|
+
|
|
21
|
+
</div>
|
|
25
22
|
|
|
26
|
-
|
|
23
|
+
If you use DeepSeek or another text-only model in DeepSeek Harness (DSH), you may have run into the same problems: the model cannot see a screenshot, generic descriptions miss the point, buttons have no usable coordinates, and a rebuilt page can look “close enough” without a way to measure the remaining difference.
|
|
27
24
|
|
|
28
|
-
|
|
25
|
+
🏆 This project is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem: it was initiated before internal beta and built during the beta with reference to [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit).
|
|
29
26
|
|
|
30
|
-
|
|
27
|
+
> **Original work:** This vision toolkit and the `vision-tools` Skill were personally created and continuously refined by the author through long-term real-world use and repeated iteration.
|
|
28
|
+
|
|
29
|
+
## Highlights
|
|
30
|
+
|
|
31
|
+
- **Paste and use it immediately.** Paste an image in DSH Web and the text-only route switches to its `(Vision Toolkit)` variant without manual path copying or model changes.
|
|
32
|
+
- **A seamless image workflow.** Native thumbnails, session history, and workspace paths stay intact; Web can preview artifacts and Headless can continue using the same structured results.
|
|
33
|
+
- **One command to install.** The built-in free Groq Qwen3.6 vision service is ready after installation, with no API key required.
|
|
34
|
+
- **Built-in free quota.** The shared service includes 100 requests per client per day, 3,000 requests globally per day, and a 60-request burst per 60 seconds, with readable errors when a limit is reached.
|
|
35
|
+
- **Vision guided by intent.** The agent extracts evidence for the task at hand, such as “Where is the error?” or “Where is the button?”, instead of returning a generic caption.
|
|
36
|
+
- **A complete screenshot-to-verification loop.** Reference images, HTML screenshots, difference regions, and pixel comparison work together for UI restoration.
|
|
37
|
+
|
|
38
|
+
[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) gives an agent more than image captions: it can read, locate, crop, trace, rebuild, and verify visual work. DSH Vision Toolkit is its native DeepSeek Harness integration, bringing that workflow into Web and Headless Profiles.
|
|
39
|
+
|
|
40
|
+
This project has two layers:
|
|
41
|
+
|
|
42
|
+
1. **Visual tools and a Skill:** the agent learns when to inspect, ground, OCR, crop, trace, or compare pixels.
|
|
43
|
+
2. **Native DSH integration:** those capabilities live inside Profiles, sessions, Settings, Artifacts, and the Web UI, with a free Groq Qwen3.6 vision service ready after installation.
|
|
44
|
+
|
|
45
|
+
> **Install and use it immediately.** The default setup includes a free Qwen3.6 vision service and requires no API key. Cropping, pixel diffing, color analysis, foreground extraction, SVG tracing, and HTML screenshots run locally without spending vision API requests.
|
|
31
46
|
|
|
32
47
|
```sh
|
|
33
48
|
dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
34
49
|
```
|
|
35
50
|
|
|
36
|
-
The npm package includes the visual toolkit snapshot and uses a managed runtime by default. **Normal installation does not require a source checkout or an `agentVisionToolkitPath`.**
|
|
37
|
-
|
|
38
51
|
**Upstream toolkit:** [Anionex/agent-vision-toolkit](https://github.com/Anionex/agent-vision-toolkit) · **Project website:** [agent-vision.anionex.me](https://agent-vision.anionex.me)
|
|
39
52
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
53
|
+
<details>
|
|
54
|
+
<summary><strong>Table of contents</strong></summary>
|
|
55
|
+
|
|
56
|
+
- [Highlights](#highlights)
|
|
57
|
+
- [Recent updates](#recent-updates)
|
|
58
|
+
- [Who it is for](#who-it-is-for)
|
|
59
|
+
- [See it in action](#see-it-in-action)
|
|
60
|
+
- [Quick start: three steps](#quick-start-three-steps)
|
|
61
|
+
- [Common workflows](#common-workflows)
|
|
62
|
+
- [Toolbox](#toolbox)
|
|
63
|
+
- [Configuration and limits](#configuration-and-limits)
|
|
64
|
+
- [Troubleshooting](#troubleshooting)
|
|
65
|
+
- [Development and community](#development-and-community)
|
|
50
66
|
|
|
51
|
-
|
|
67
|
+
</details>
|
|
52
68
|
|
|
53
|
-
##
|
|
69
|
+
## Recent updates
|
|
54
70
|
|
|
55
|
-
|
|
71
|
+
- **2026-08-16 · Windows Python:** Added Microsoft Store Python support, fixing first-time isolated-runtime setup failures for affected Windows users.
|
|
72
|
+
- **2026-08-16 · Better free vision:** Switched the built-in no-key service to Groq Qwen3.6, improving image understanding without adding setup steps.
|
|
73
|
+
- **2026-08-16 · Image paste:** Text-only routes now switch to a `(Vision Toolkit)` variant and keep a workspace path, fixing blocked pastes and images that could not be reused later.
|
|
74
|
+
- **2026-08-16 · Higher free quotas:** Raised the shared service ceiling to `3,000/day` and `60/minute` to make better use of the three-account Groq pool while keeping the per-client limit at `100/day`.
|
|
75
|
+
- **2026-08-16 · Real model test:** Added a full image-request test in Settings, fixing the false confidence caused by a successful `/models` request to a model that still cannot process images.
|
|
56
76
|
|
|
57
|
-
|
|
77
|
+
## Who it is for
|
|
58
78
|
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
79
|
+
| The problem | What Vision Toolkit delivers |
|
|
80
|
+
|---|---|
|
|
81
|
+
| **A text-only model cannot see a screenshot** | Paste an image in DSH Web; the plugin obtains visual evidence and returns the task-relevant parts to the text model |
|
|
82
|
+
| **The description is long but misses the point** | Ask “Where is the error?” or “What color is the submit button?” and receive an answer focused on that question |
|
|
83
|
+
| **The model knows an element exists but cannot act on it** | Get original-image pixel coordinates and an optional labeled or numbered preview |
|
|
84
|
+
| **Long-screenshot OCR skips or duplicates lines** | Split and audit the image while preserving Markdown, chunks, manifests, and resumable run state |
|
|
85
|
+
| **UI restoration is judged by feel** | Compare the reference and implementation screenshots to get a difference percentage, ranked regions, a heatmap, and JSON |
|
|
86
|
+
| **Screenshot assets cannot be reused** | Produce a crop, transparent PNG, color palette, or editable SVG instead of stopping at prose |
|
|
62
87
|
|
|
63
|
-
|
|
88
|
+
## See it in action
|
|
64
89
|
|
|
65
|
-
###
|
|
90
|
+
### Paste an image directly into DSH
|
|
66
91
|
|
|
67
92
|
<p align="center">
|
|
68
|
-
<img src="assets/
|
|
69
|
-
<img src="assets/upstream/infographic-result.webp" width="49%" alt="Upstream editable HTML and CSS reconstruction of the model-training infographic." />
|
|
93
|
+
<img src="assets/dsh-view-example.png" width="82%" alt="A text-only DeepSeek model answering a question about a pasted image through Vision Toolkit in DSH Web" />
|
|
70
94
|
</p>
|
|
71
95
|
|
|
72
|
-
*
|
|
96
|
+
*Paste an image into the conversation. A text-only model can switch to its `Vision Toolkit` variant and inspect the image in the context of the user's question.*
|
|
73
97
|
|
|
74
|
-
###
|
|
98
|
+
### Screenshot to editable page
|
|
75
99
|
|
|
76
100
|
<p align="center">
|
|
77
|
-
<img src="assets/upstream/
|
|
78
|
-
<img src="assets/upstream/
|
|
101
|
+
<img src="assets/upstream/infographic-reference.webp" width="49%" alt="Reference infographic screenshot used for restoration" />
|
|
102
|
+
<img src="assets/upstream/infographic-result.webp" width="49%" alt="Editable HTML and CSS reconstruction created from the reference screenshot" />
|
|
79
103
|
</p>
|
|
80
104
|
|
|
81
|
-
*Left:
|
|
105
|
+
*Left: the reference screenshot. Right: an editable HTML/CSS result. The result can continue into screenshot rendering and pixel comparison instead of ending as an image description.*
|
|
82
106
|
|
|
83
|
-
###
|
|
107
|
+
### Sketch to working interface
|
|
84
108
|
|
|
85
109
|
<p align="center">
|
|
86
|
-
<img src="assets/
|
|
87
|
-
<img src="assets/
|
|
110
|
+
<img src="assets/upstream/ui-sketch.webp" width="49%" alt="Hand-drawn JupyterLab interface used as the restoration reference" />
|
|
111
|
+
<img src="assets/upstream/ui-result.webp" width="49%" alt="Working JupyterLab-style interface reconstructed from the sketch" />
|
|
88
112
|
</p>
|
|
89
113
|
|
|
90
|
-
*Left:
|
|
114
|
+
*Left: a hand-drawn reference. Right: the working interface reconstructed from it.*
|
|
91
115
|
|
|
92
|
-
|
|
116
|
+
### Turn “looks close” into a verifiable result
|
|
93
117
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
The included UI-restoration workflow starts with an intentionally inaccurate HTML implementation. Vision Toolkit measures a `6.04%` difference, points to the worst regions, and helps drive the next iteration. The final render reaches an exact `0%` difference at `1200 × 720`.
|
|
118
|
+
The repository includes a reproducible UI-restoration example: the agent renders the reference and implementation, then uses difference regions, a heatmap, and a JSON report to guide the next correction.
|
|
97
119
|
|
|
98
120
|
<p>
|
|
99
|
-
<img src="examples/ui-restoration/assets/initial.png" width="49%" alt="Initial UI
|
|
100
|
-
<img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="
|
|
121
|
+
<img src="examples/ui-restoration/assets/initial.png" width="49%" alt="Initial UI implementation with measurable layout and styling differences" />
|
|
122
|
+
<img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="UI implementation after visual diagnosis and pixel comparison" />
|
|
101
123
|
</p>
|
|
102
124
|
|
|
103
|
-
|
|
104
|
-
|---|---|
|
|
105
|
-
| Reference image | A working HTML implementation you can open and edit |
|
|
106
|
-
| First comparison | `6.04%` difference across the visible problem regions |
|
|
107
|
-
| Final comparison | `0%` difference at `1200 × 720` |
|
|
108
|
-
|
|
109
|
-
## Why it feels different
|
|
110
|
-
|
|
111
|
-
- **Ask for the thing you need.** “Where is the submit button?” and “Why does this screenshot differ from the reference?” lead to focused visual work instead of a generic caption.
|
|
112
|
-
- **Get evidence you can use.** The agent returns coordinates, OCR, measurements, JSON, and files you can open or pass to the next step.
|
|
113
|
-
- **Keep the workflow in DSH.** Credentials, Settings, Artifacts, Web cards, and Headless results live alongside the rest of your session.
|
|
114
|
-
- **Use local tools when you can.** Crop, trace, pixel comparison, color analysis, foreground extraction, and HTML screenshots do not consume a vision API request.
|
|
115
|
-
- **Repeat the loop.** Reference image → implementation → screenshot → pixel diff gives UI work a measurable finish line.
|
|
125
|
+
## Quick start: three steps
|
|
116
126
|
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
Use DeepSeek Harness `0.1.0-rc.6` or a compatible later `0.1.x` release. The plugin prepares its managed runtime on first use.
|
|
127
|
+
### 1. Install
|
|
120
128
|
|
|
121
129
|
```sh
|
|
122
130
|
dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
123
|
-
dsh plugin --profile headless add @anionex/dsh-vision-toolkit
|
|
124
131
|
```
|
|
125
132
|
|
|
126
|
-
|
|
127
|
-
2. New installations use the built-in free Gemma 4 provider, so you can run **Test API connection** and **Test vision model** without an API key. To use another provider, edit the endpoint/model/protocol and provide its DSH Credential.
|
|
128
|
-
3. In a conversation, paste an image or put it in the workspace, invoke `/vision-tools`, and ask for a concrete visual task.
|
|
129
|
-
|
|
130
|
-
If you use an older DSH launcher, the profile may need `nodeLinker: hoisted` and `autoInstallPeers: false` before installation. Current launchers repair these settings for you.
|
|
131
|
-
|
|
132
|
-
Local crop, trace, pixel, color, foreground, and HTML operations do not require a visual API credential.
|
|
133
|
+
You can install it into a Headless Profile too:
|
|
133
134
|
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
Join the `agent-vision-toolkit` community group to exchange usage tips, share feedback, and suggest improvements.
|
|
137
|
-
|
|
138
|
-
<p align="center">
|
|
139
|
-
<img src="assets/community-group-qr.png" alt="QR code for the agent-vision-toolkit community group" width="260">
|
|
140
|
-
</p>
|
|
141
|
-
|
|
142
|
-
> **No local path is required.** Keep the default `runtime.mode: managed` for the normal npm installation. The optional `runtime.agentVisionToolkitPath` setting is only for developers or controlled deployments that deliberately use an external pinned checkout.
|
|
143
|
-
|
|
144
|
-
<details>
|
|
145
|
-
<summary><strong>Technical architecture</strong></summary>
|
|
146
|
-
|
|
147
|
-
## How it works
|
|
148
|
-
|
|
149
|
-
```mermaid
|
|
150
|
-
flowchart LR
|
|
151
|
-
User["Workspace image or local HTML"] --> Skill["vision-tools Skill"]
|
|
152
|
-
Skill --> Activate["Agent-scoped activation"]
|
|
153
|
-
Activate --> Tools["10 independent vision_* tools"]
|
|
154
|
-
Tools --> Runtime["Shared VisionToolkitRuntime"]
|
|
155
|
-
Credentials["DSH Credentials"] --> Runtime
|
|
156
|
-
Settings["Web Settings and health"] --> Runtime
|
|
157
|
-
Runtime --> Upstream["Pinned agent-vision-toolkit"]
|
|
158
|
-
Runtime --> Remote["Configured vision API"]
|
|
159
|
-
Upstream --> Result["Text, coordinates, JSON"]
|
|
160
|
-
Remote --> Result
|
|
161
|
-
Runtime --> Artifacts["Workspace Artifacts"]
|
|
162
|
-
Result --> Session["Reconstructable Session log"]
|
|
163
|
-
Artifacts --> Web["Preview, download, or open file"]
|
|
135
|
+
```sh
|
|
136
|
+
dsh plugin --profile headless add @anionex/dsh-vision-toolkit
|
|
164
137
|
```
|
|
165
138
|
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
</details>
|
|
139
|
+
### 2. Restart and check it
|
|
169
140
|
|
|
170
|
-
|
|
141
|
+
Restart a running Web Profile, then open **Settings → Vision Toolkit**. The free provider is already configured; run **Test vision model** to confirm it is reachable.
|
|
171
142
|
|
|
172
|
-
|
|
173
|
-
|---|---|---|---|
|
|
174
|
-
| `vision_glance` | Remote vision API | Description, targeted answer, OCR, or multi-image comparison | None |
|
|
175
|
-
| `vision_ground` | Remote vision API; optional local preview | Target, original-image dimensions, and pixel boxes | Optional labeled PNG |
|
|
176
|
-
| `vision_detect` | Remote vision API; optional local preview | Numbered element inventory and original-image pixel boxes | Optional numbered PNG |
|
|
177
|
-
| `vision_trace` | Local pinned vtracer pipeline | SVG geometry status, path count, scale, and size | SVG |
|
|
178
|
-
| `vision_crop` | Local Pillow pipeline | Applied pixel box, dimensions, format, and clamp status | PNG or JPEG |
|
|
179
|
-
| `vision_pixel_diff` | Local NumPy/Pillow pipeline | Difference percentage and ranked grid regions | PNG heatmap and JSON report |
|
|
180
|
-
| `vision_long_screenshot_ocr` | Local split/audit; remote OCR unless `splitOnly=true` | Chunk boundaries, reuse state, completion state, and run directory | Markdown, manifest, boundary audit, chunk PNGs, and OCR sidecars |
|
|
181
|
-
| `vision_extract_foreground` | Local pinned extraction pipeline | Selected box, component counts, foreground coverage, and dimensions | Transparent PNG |
|
|
182
|
-
| `vision_dominant_colors` | Local pinned color analysis | Extracted palette or pixel-backed candidate ranking | None |
|
|
183
|
-
| `vision_html_screenshot` | Local Chrome/Chromium/Edge adapter | Authorized source facts, viewport, and rendered dimensions | PNG |
|
|
143
|
+
The first start prepares an isolated runtime, so it needs access to the Python package cache or the network. A normal installation does not require an `agent-vision-toolkit` source checkout or a local path setting.
|
|
184
144
|
|
|
185
|
-
|
|
145
|
+
### 3. Paste an image and describe the outcome you want
|
|
186
146
|
|
|
187
|
-
|
|
188
|
-
<summary><strong>Advanced model behavior</strong></summary>
|
|
147
|
+
Paste a screenshot into the conversation or place an image in the session workspace, then invoke `/vision-tools`. For example:
|
|
189
148
|
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
## Image-input variants for text-only models
|
|
197
|
-
|
|
198
|
-
Text-only model routes get sibling model-selector entries named `<model> (Vision Toolkit)` under a matching provider group. DSH cannot pass a pasted attachment's local path through its native image block, so every bridge path materializes the image inside the session workspace and exposes its absolute path to the model. The model can then call `vision_glance` (or another visual tool) with that path. When the server-side image-input variant is active, the same model-visible block also contains the focus-hinted `[vision model description]` evidence aligned with `agent-vision-toolkit`; the path remains available for a second, more targeted visual call. The session log contains the durable path reference and the UI keeps the paste record.
|
|
199
|
-
|
|
200
|
-
A variant is registered automatically for every model the host positively declares text-only (for example the DeepSeek chat family). With the default `autoSwitch: true`, the browser switches to `<model> (Vision Toolkit)` and the server-side bridge rewrites each native image block into **both** the workspace path and the focus-hinted description; the path is not hidden from the model. Setting `autoSwitch: false` keeps the older path-only takeover instead. The host's verdict uses the exact model route the browser read from the live model catalog, with the selector label as fallback; unconfirmed or image-capable routes keep their native flow.
|
|
201
|
-
|
|
202
|
-
Description conversion needs the configured vision provider and its credential when the opt-in image-input variant is used; when the runtime is not ready or a read fails, the wire block keeps the workspace path and adds the upstream-compatible `[vision unavailable: ...]` note instead of failing the turn. The bridge does not treat injected context files as the current user intent, and it uses the latest assistant paragraph when a tool-fetched image is being described. Disable variants with `imageInputVariants.enabled: false`, restrict the wrapped routes with `imageInputVariants.providers`, or opt into native attachment switching with `imageInputVariants.autoSwitch: true`.
|
|
203
|
-
|
|
204
|
-
</details>
|
|
205
|
-
|
|
206
|
-
## Requirements
|
|
207
|
-
|
|
208
|
-
- DeepSeek Harness with a Web or Headless profile and `pnpm` available to `dsh plugin`.
|
|
209
|
-
- Python 3.11 or newer. Managed mode creates an isolated environment, so users do not install the upstream CLI or Python packages manually.
|
|
210
|
-
- Network access on the first managed-runtime activation unless the exact packages in `runtime/requirements.lock` are already available in the configured package cache.
|
|
211
|
-
- The built-in free Gemma 4 provider is ready for `vision_glance`, `vision_ground`, `vision_detect`, and non-split-only long-screenshot OCR. A DSH Credential is required only when a custom OpenAI-compatible or Anthropic endpoint is configured. Local tools remain usable without either provider.
|
|
212
|
-
- Chrome, Chromium, or Edge only for `vision_html_screenshot`; all other tools remain available when no supported browser is installed.
|
|
213
|
-
- PNG, JPEG, GIF, or WebP inputs inside the session workspace or an explicitly configured `allowedDirs` root.
|
|
214
|
-
|
|
215
|
-
## Install and lifecycle
|
|
216
|
-
|
|
217
|
-
### Install
|
|
218
|
-
|
|
219
|
-
Install the bundle into each profile that should expose it:
|
|
220
|
-
|
|
221
|
-
```sh
|
|
222
|
-
dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
223
|
-
dsh plugin --profile headless add @anionex/dsh-vision-toolkit
|
|
224
|
-
dsh --profile web --dump-config | grep vision-toolkit
|
|
225
|
-
dsh --profile headless --dump-config | grep vision-toolkit
|
|
149
|
+
```text
|
|
150
|
+
Inspect this screenshot. Explain the error and tell me what to fix first.
|
|
151
|
+
Find the login button in the top-right corner, return original pixel coordinates, and make a boxed preview.
|
|
152
|
+
Crop this icon and convert it to SVG.
|
|
153
|
+
Rebuild the page from reference.png. After each pass, render it and run a pixel diff until the major differences are gone.
|
|
226
154
|
```
|
|
227
155
|
|
|
228
|
-
|
|
156
|
+
## Common workflows
|
|
229
157
|
|
|
230
|
-
|
|
158
|
+
| Task | Recommended workflow |
|
|
159
|
+
|---|---|
|
|
160
|
+
| Image Q&A or screenshot debugging | Inspect → answer around the current question → locate details when needed |
|
|
161
|
+
| Find a button, icon, or text region | Ground the target → return pixel box → create a labeled preview |
|
|
162
|
+
| Extract an icon from a screenshot | Ground → crop → trace to SVG |
|
|
163
|
+
| Read a long webpage screenshot | Split → OCR → merge Markdown → audit boundaries |
|
|
164
|
+
| Recreate a page or component | Reference → implementation → HTML screenshot → pixel diff → iterate |
|
|
165
|
+
| Extract brand visuals | Crop region → analyze dominant colors → extract foreground → export transparent PNG |
|
|
231
166
|
|
|
232
|
-
|
|
167
|
+
## Toolbox
|
|
233
168
|
|
|
234
|
-
|
|
169
|
+
The plugin provides 10 tools that can be called independently or composed into a workflow:
|
|
235
170
|
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
|
|
171
|
+
| Tool | Best question to ask | Main result |
|
|
172
|
+
|---|---|---|
|
|
173
|
+
| `vision_glance` | “What is happening in this image?” | Focused answer, description, OCR, or multi-image comparison |
|
|
174
|
+
| `vision_ground` | “Where is the thing I need?” | Original pixel coordinates and optional boxed preview |
|
|
175
|
+
| `vision_detect` | “Which buttons, icons, or elements are present?” | Numbered element inventory, coordinates, and optional preview |
|
|
176
|
+
| `vision_crop` | “Extract this region as its own image” | PNG or JPEG crop |
|
|
177
|
+
| `vision_trace` | “Turn this shape into an editable vector” | SVG |
|
|
178
|
+
| `vision_pixel_diff` | “Where does the implementation differ from the reference?” | Difference percentage, ranked regions, heatmap, and JSON |
|
|
179
|
+
| `vision_long_screenshot_ocr` | “Read this entire long screenshot” | Markdown, chunks, manifest, and audit output |
|
|
180
|
+
| `vision_extract_foreground` | “Remove the background from this subject” | Transparent PNG |
|
|
181
|
+
| `vision_dominant_colors` | “Which colors dominate this area?” | Palette or ranked candidate colors |
|
|
182
|
+
| `vision_html_screenshot` | “Render this local page at an exact viewport or capture the full page” | PNG and optional CSS `pageHeight` |
|
|
242
183
|
|
|
243
|
-
|
|
184
|
+
Coordinates always use original-image pixels in `x1,y1,x2,y2` form, so grounding output can feed directly into cropping, tracing, or later automation.
|
|
244
185
|
|
|
245
|
-
|
|
186
|
+
For a long HTML document, pass `fullPage=true`. The requested width and height remain the layout viewport, while the resulting PNG covers the complete document and reports `pageHeight` in CSS pixels.
|
|
246
187
|
|
|
247
|
-
|
|
248
|
-
dsh plugin --profile web remove @dsh-external/dsh-vision-toolkit
|
|
249
|
-
dsh plugin --profile web add @anionex/dsh-vision-toolkit
|
|
250
|
-
```
|
|
188
|
+
## How it works
|
|
251
189
|
|
|
252
|
-
|
|
190
|
+
The plugin keeps image understanding and deterministic local image processing in one Agent workflow. Expand the flow below for the implementation boundary.
|
|
253
191
|
|
|
254
|
-
|
|
192
|
+
<details>
|
|
193
|
+
<summary><strong>Architecture and image-input behavior</strong></summary>
|
|
255
194
|
|
|
256
|
-
```
|
|
257
|
-
|
|
258
|
-
|
|
195
|
+
```mermaid
|
|
196
|
+
flowchart LR
|
|
197
|
+
Image["Screenshot or local HTML"] --> Skill["vision-tools Skill"]
|
|
198
|
+
Skill --> Agent["Text agent selects a task"]
|
|
199
|
+
Agent --> Vision["Use a vision model when image understanding is needed"]
|
|
200
|
+
Agent --> Local["Run crop, SVG, and pixel work locally"]
|
|
201
|
+
Vision --> Result["Answer, OCR, coordinates"]
|
|
202
|
+
Local --> Artifact["PNG, SVG, heatmap, JSON"]
|
|
203
|
+
Result --> Session["Continue reasoning and acting"]
|
|
204
|
+
Artifact --> Session
|
|
259
205
|
```
|
|
260
206
|
|
|
261
|
-
|
|
207
|
+
The visual capabilities come from a packaged, pinned `agent-vision-toolkit` snapshot. The DSH plugin handles installation, session-scoped tool exposure, Credentials, path checks, cancellation, timeouts, result files, and Web presentation. The runtime never fetches upstream `main` in the background.
|
|
262
208
|
|
|
263
|
-
|
|
209
|
+
For routes that DSH positively identifies as text-only, the plugin registers a sibling `<model> (Vision Toolkit)` variant. By default, pasting an image in DSH Web switches to that variant and gives the model both a reusable workspace path and a visual description focused on the current task.
|
|
264
210
|
|
|
265
|
-
|
|
266
|
-
dsh plugin --profile web remove @anionex/dsh-vision-toolkit
|
|
267
|
-
dsh plugin --profile headless remove @anionex/dsh-vision-toolkit
|
|
268
|
-
```
|
|
211
|
+
</details>
|
|
269
212
|
|
|
270
|
-
|
|
213
|
+
## Configuration and limits
|
|
271
214
|
|
|
272
|
-
|
|
215
|
+
### Built-in free service
|
|
273
216
|
|
|
274
|
-
The
|
|
217
|
+
The default setup uses:
|
|
275
218
|
|
|
276
|
-
```
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
baseUrl: https://vision.anionex.me/v1
|
|
281
|
-
credential: ANIONEX_FREE_VISION
|
|
282
|
-
model: gemma-4-26b-a4b-it
|
|
283
|
-
protocol: openai
|
|
284
|
-
anthropicThinking: omit
|
|
285
|
-
userAgent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36
|
|
286
|
-
language: zh
|
|
287
|
-
timeoutMs: 60000
|
|
288
|
-
maxImageBytes: 4194304
|
|
289
|
-
maxImagePixels: 20000000
|
|
290
|
-
concurrency: 4
|
|
291
|
-
runtime:
|
|
292
|
-
mode: managed
|
|
293
|
-
allowedDirs: []
|
|
294
|
-
imageInputVariants:
|
|
295
|
-
enabled: true
|
|
296
|
-
providers: []
|
|
297
|
-
autoSwitch: true
|
|
219
|
+
```text
|
|
220
|
+
Base URL: https://vision.anionex.me/v1
|
|
221
|
+
Model: qwen/qwen3.6-27b
|
|
222
|
+
API Key: no user configuration required
|
|
298
223
|
```
|
|
299
224
|
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
| Field | Default | Contract |
|
|
303
|
-
|---|---|---|
|
|
304
|
-
| `provider.baseUrl` | `https://vision.anionex.me/v1` | Built-in free OpenAI-compatible endpoint; custom providers may use another base URL, normalized without trailing slashes |
|
|
305
|
-
| `provider.credential` | `ANIONEX_FREE_VISION` | Read-only built-in reference for the free service; custom providers use a DSH Credential reference, never a secret value |
|
|
306
|
-
| `provider.model` | `gemma-4-26b-a4b-it` | Multimodal model name sent to remote tools |
|
|
307
|
-
| `provider.protocol` | `openai` | `openai` sends Chat Completions requests; `anthropic` sends native Messages requests |
|
|
308
|
-
| `provider.anthropicThinking` | `omit` | Anthropic thinking field. `omit` sends no thinking field and has the broadest compatibility. Use `disabled` or `adaptive` only when the selected model documents that mode; restore `omit` first if the provider returns HTTP 400. |
|
|
309
|
-
| `provider.userAgent` | browser-compatible default | User-Agent sent by vision requests and explicit connection tests; override it for provider or proxy compatibility |
|
|
310
|
-
| `language` | `zh` | Vision output language: `zh` or `en` |
|
|
311
|
-
| `timeoutMs` | `60000` | Whole-operation deadline, 1000-600000 ms; each tool may request a narrower override |
|
|
312
|
-
| `maxImageBytes` | `4194304` | Encoded-byte limit per input image; the built-in free service accepts up to 4 MiB |
|
|
313
|
-
| `maxImagePixels` | `20000000` | Decoded-pixel limit per input image; the built-in free service accepts up to 20,000,000 pixels |
|
|
314
|
-
| `concurrency` | `4` | In-flight operations per session, 1-16 |
|
|
315
|
-
| `runtime.mode` | `managed` | `managed` uses the packaged snapshot; `external` accepts only the exact pin |
|
|
316
|
-
| `runtime.agentVisionToolkitPath` | unset | Required in `external` mode; exported exact snapshot or clean pinned Git checkout |
|
|
317
|
-
| `runtime.python` | unset | Optional Python 3.11+ bootstrap/interpreter override |
|
|
318
|
-
| `allowedDirs` | `[]` | Additional realpath-resolved input roots; the session workspace is always allowed |
|
|
319
|
-
| `imageInputVariants.enabled` | `true` | Register image-input variant entries for text-only model routes in the model selector |
|
|
320
|
-
| `imageInputVariants.providers` | `[]` | Restrict wrapped upstream routes by provider id; empty wraps every eligible route |
|
|
321
|
-
| `imageInputVariants.autoSwitch` | `true` | Automatically switch a text-only session to its image-input variant; the model receives both the workspace path and focused description. `false` keeps the path-only takeover |
|
|
322
|
-
|
|
323
|
-
### Credentials
|
|
324
|
-
|
|
325
|
-
The built-in free provider uses the fixed `ANIONEX_FREE_VISION` reference and does not accept or store a user API key. If you change the endpoint, model, or protocol to a custom provider, the write-only **API key** field unlocks; saving a non-empty value writes it under the advanced **Credential name** reference. Headless deployments can pre-provision that custom reference in `$DSH_HOME/.credentials.yaml`.
|
|
326
|
-
|
|
327
|
-
Settings store only the reference, never the value. The browser does not receive a stored value, and a successful save clears the field instead of echoing it. Remote operations resolve the reference once per call and inject the value only into that subprocess environment. The plugin excludes user `.env` files, checkout `.env` files, `PYTHONPATH`, `PYTHONHOME`, `VIRTUAL_ENV`, and user site-packages so ambient Python or upstream configuration cannot override the selected DSH provider. Logs, errors, tool results, Artifact metadata, and Settings responses never contain the secret.
|
|
328
|
-
|
|
329
|
-
### Built-in free service limits
|
|
330
|
-
|
|
331
|
-
The public service is shared and intended as a zero-configuration default, not an unlimited private endpoint. Limits are enforced by the proxy and returned as OpenAI-style errors with a reason code and readable message; rate-limit responses also include `Retry-After` and request-quota headers.
|
|
225
|
+
This is a shared zero-configuration entry point, not an unlimited private endpoint. Current limits are:
|
|
332
226
|
|
|
333
227
|
| Limit | Current value |
|
|
334
228
|
|---|---:|
|
|
335
229
|
| Per client | 100 requests per UTC day |
|
|
336
|
-
|
|
|
337
|
-
| Burst |
|
|
338
|
-
| Image
|
|
230
|
+
| Whole service | 3,000 requests per UTC day |
|
|
231
|
+
| Burst | 60 requests per 60 seconds |
|
|
232
|
+
| Image size | 4 MiB per image |
|
|
339
233
|
| Decoded pixels | 20,000,000 per image |
|
|
340
|
-
| Output | 512 tokens
|
|
234
|
+
| Output | 512 tokens per request |
|
|
341
235
|
|
|
342
|
-
|
|
236
|
+
The limits protect shared capacity and prevent unusually large images from monopolizing memory or request time. When a limit is reached, the service returns a readable reason and error code. Rate-limit responses also include `Retry-After` instead of collapsing into an unexplained model failure.
|
|
343
237
|
|
|
344
|
-
|
|
238
|
+
### Bring your own vision model
|
|
345
239
|
|
|
346
|
-
|
|
240
|
+
For higher quotas, private endpoints, or another model, change the provider in **Settings → Vision Toolkit** and store the API key as a DSH Credential. Settings stores the Credential reference and never reads the saved secret back into the browser.
|
|
347
241
|
|
|
348
|
-
|
|
242
|
+
You can also configure a Profile patch:
|
|
349
243
|
|
|
350
244
|
```yaml
|
|
351
245
|
- id: vision-toolkit
|
|
352
246
|
config:
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
247
|
+
provider:
|
|
248
|
+
baseUrl: https://api.example.com/v1
|
|
249
|
+
credential: MY_VISION_KEY
|
|
250
|
+
model: your-vision-model
|
|
251
|
+
protocol: openai
|
|
357
252
|
```
|
|
358
253
|
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
## Web Settings
|
|
254
|
+
OpenAI Chat Completions-compatible endpoints and Anthropic Messages are supported. The Web Settings panel exposes the full provider, runtime, timeout, image-limit, and image-input-variant configuration.
|
|
362
255
|
|
|
363
|
-
|
|
256
|
+
### Requirements
|
|
364
257
|
|
|
365
|
-
|
|
258
|
+
- A DeepSeek Harness Web or Headless Profile.
|
|
259
|
+
- Node.js `^22.19.0` or `>=24.0.0`.
|
|
260
|
+
- Python 3.11+; the plugin prepares an isolated environment by default.
|
|
261
|
+
- Only `vision_html_screenshot` requires Chrome, Chromium, or Edge.
|
|
262
|
+
- Inputs must be PNG, JPEG, GIF, or WebP files in the session workspace or an explicitly allowed directory.
|
|
366
263
|
|
|
367
|
-
|
|
368
|
-
|
|
369
|
-
`Run health check` performs local checks only. `Test API connection` is an explicit action that sends the configured Credential to `GET /models`; OpenAI uses Bearer authentication, while Anthropic uses `x-api-key` and `anthropic-version`. That lightweight probe uploads no image and creates no completion. `Test vision model` separately sends the bundled `assets/vision-model-test.png` through the same multimodal runtime path as `vision_glance`; it creates one real completion and is the authoritative check that the selected endpoint, credential, model, protocol, and upstream account can process images. The Vision model health card displays a dedicated `Verified`, `Not tested`, or `Test failed` tag, so an HTTP 200 response from `/models` is not presented as a successful image test. Plugin load and ordinary Settings reads never make either request.
|
|
370
|
-
|
|
371
|
-
Health, connection testing, and plugin/upstream version inspection are administrative Web Settings capabilities rather than model-facing tools, so their schemas never occupy an agent request.
|
|
372
|
-
|
|
373
|
-
## Artifacts and presentation
|
|
374
|
-
|
|
375
|
-
Artifact-producing tools write only under `<workspace>/.dsh-vision-toolkit/artifacts`, either as one validated file or an atomically committed run directory. Each model-visible descriptor contains the path, filename, MIME type, kind, description, source tool, preview intent, and byte size, so Headless agents can reuse the path in later calls without browser support. Before a traced SVG is committed, the runtime parses it as XML: standard declarations and comments are accepted, while doctypes, malformed or multi-root documents, a non-SVG namespace, and reported path/byte mismatches are rejected.
|
|
264
|
+
<details>
|
|
265
|
+
<summary><strong>Install, upgrade, disable, and uninstall</strong></summary>
|
|
376
266
|
|
|
377
|
-
|
|
267
|
+
```sh
|
|
268
|
+
dsh plugin --profile web update @anionex/dsh-vision-toolkit
|
|
269
|
+
dsh plugin --profile web remove @anionex/dsh-vision-toolkit
|
|
270
|
+
```
|
|
378
271
|
|
|
379
|
-
|
|
272
|
+
If you are migrating from the retired `@dsh-external/dsh-vision-toolkit`, remove the old package first and install `@anionex/dsh-vision-toolkit`.
|
|
380
273
|
|
|
381
|
-
|
|
274
|
+
To disable the bundle temporarily, set this in the Profile patch:
|
|
382
275
|
|
|
383
|
-
```
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
vision_detect image="screenshot.png" category="buttons" preview=true
|
|
387
|
-
vision_crop image="screenshot.png" region="1067,841,1108,881"
|
|
388
|
-
vision_trace image="icon.png" color=true output="icon.svg"
|
|
389
|
-
vision_pixel_diff original="reference.png" rebuilt="actual.png" runName="comparison"
|
|
390
|
-
vision_long_screenshot_ocr image="page.png" mode="general" jobs=2
|
|
391
|
-
vision_extract_foreground image="logo.png" mode="color"
|
|
392
|
-
vision_dominant_colors image="screen.png" region="0,0,600,300" top=8
|
|
393
|
-
vision_html_screenshot source="implementation.html" width=1200 height=720
|
|
276
|
+
```yaml
|
|
277
|
+
- id: vision-toolkit
|
|
278
|
+
disabled: true
|
|
394
279
|
```
|
|
395
280
|
|
|
396
|
-
|
|
281
|
+
Restart the Web Profile and refresh the page after enabling or upgrading the Web plugin.
|
|
397
282
|
|
|
398
|
-
|
|
283
|
+
</details>
|
|
399
284
|
|
|
400
|
-
|
|
285
|
+
### Plugin updates
|
|
401
286
|
|
|
402
|
-
|
|
403
|
-
npm run example:ui-restoration
|
|
404
|
-
npm run example:ui-restoration:write
|
|
405
|
-
```
|
|
287
|
+
In **Settings → Vision Toolkit**, **Check for updates** queries the Profile's npm registry. For a direct registry installation, **Update and restart** installs only the exact version you confirmed, verifies it, and restarts an explicitly opted-in POSIX Web process on a fixed `--port`. Local/workspace/file/git/URL installs, Windows, dynamic ports, read-only Profiles, and manager-owned processes remain check-only.
|
|
406
288
|
|
|
407
|
-
The
|
|
289
|
+
The updater revalidates the Profile before mutation, snapshots the original manifest and lockfile, and holds a token-owned cross-process lock. The current Web process exits only after the restart helper confirms that the backup is readable and the lock handoff succeeded. When the Profile was already operational, the replacement must report both the target plugin version and a ready runtime; failed replacements restore the original manifest/lockfile and rebuild dependencies with a frozen lockfile before retrying the previous exact version. If automatic recovery itself fails, the backup and lock are preserved and their paths are written to `$DSH_HOME/logs/vision-toolkit-restart.log`. Detached restart requires `DSH_VISION_TOOLKIT_ALLOW_DETACHED_RESTART=1`; unsaved Settings or API-key input blocks installation.
|
|
408
290
|
|
|
409
291
|
## Troubleshooting
|
|
410
292
|
|
|
411
|
-
|
|
|
293
|
+
| Problem | What to do |
|
|
412
294
|
|---|---|
|
|
413
|
-
|
|
|
414
|
-
|
|
|
415
|
-
|
|
|
416
|
-
|
|
|
417
|
-
|
|
|
418
|
-
|
|
|
419
|
-
|
|
|
420
|
-
|
|
421
|
-
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
## Development and verification
|
|
295
|
+
| Pasting an image still says the model does not support image input | Restart the Web Profile, refresh the page, and confirm the selected route has the `(Vision Toolkit)` suffix. You can also place the image in the session workspace and invoke `/vision-tools` |
|
|
296
|
+
| The free service returns 429 | Wait for the `Retry-After` interval, or switch to your own endpoint when you need stable higher volume |
|
|
297
|
+
| The image exceeds a size or pixel limit | Crop or resize it first; the error identifies whether bytes or decoded pixels caused the rejection |
|
|
298
|
+
| A custom Credential is missing | Enter the API key in **Settings → Vision Toolkit** and confirm the Credential name matches the provider configuration |
|
|
299
|
+
| First-time runtime setup fails | Check Python 3.11+, network or package-cache access, and disk permissions, then retry the model test in Settings |
|
|
300
|
+
| Chrome is not found | Install Chrome, Chromium, or Edge. Only HTML screenshot rendering is unavailable; the other tools still work |
|
|
301
|
+
| An artifact cannot be previewed | Use **Open file** or the workspace path in the result. Preview URLs exist only while the Web route is available |
|
|
302
|
+
|
|
303
|
+
## Project status and limitations
|
|
304
|
+
|
|
305
|
+
The current release focuses on screenshot understanding, visual grounding, OCR, asset extraction, UI restoration, and pixel-level verification. It is not a video, audio, or camera-input system and does not automatically click GUI controls. Interactive box editing, remote service clusters, model voting, and cross-session visual caches are also outside the current scope.
|
|
306
|
+
|
|
307
|
+
## Development and community
|
|
427
308
|
|
|
428
309
|
```sh
|
|
429
310
|
pnpm install --frozen-lockfile --trust-lockfile
|
|
430
311
|
pnpm run verify:portable
|
|
431
312
|
pnpm run build
|
|
432
313
|
pnpm test
|
|
433
|
-
pnpm
|
|
434
|
-
pnpm pack --dry-run
|
|
314
|
+
TSX_TSCONFIG_PATH=tsconfig.json pnpm dlx tsx scripts/ui-restoration-example.ts --check
|
|
435
315
|
```
|
|
436
316
|
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
## Project status and scope
|
|
444
|
-
|
|
445
|
-
Version `0.1.12` is the current public npm release. The product focuses on screenshot understanding, visual grounding, OCR, asset extraction, UI restoration, and pixel-level verification in DSH Web and Headless profiles. Web upload, drag-and-drop, camera/video/audio/document ingestion, interactive box editing, automatic GUI clicking, service clusters, model routing, model voting, and cross-session vision caches remain outside the current product.
|
|
446
|
-
|
|
447
|
-
<details>
|
|
448
|
-
<summary><strong>Maintainer scope note</strong></summary>
|
|
449
|
-
|
|
450
|
-
The stable `ctx.visionToolkit` service and capability-discovery API remain unpublished until an independent plugin becomes a real consumer. This keeps the public integration surface tied to a tested use case rather than an unvalidated ecosystem contract.
|
|
451
|
-
|
|
452
|
-
</details>
|
|
317
|
+
- Read [CONTRIBUTING.md](CONTRIBUTING.md) before contributing.
|
|
318
|
+
- Use [GitHub Issues](https://github.com/Anionex/dsh-vision-toolkit/issues) for bugs, focused feature requests, and usage questions; see [SUPPORT.md](SUPPORT.md) for channel guidance.
|
|
319
|
+
- Report vulnerabilities privately through [SECURITY.md](SECURITY.md).
|
|
320
|
+
- See [CHANGELOG.md](CHANGELOG.md) for releases and [FUNDING.md](FUNDING.md) for sponsorship details.
|
|
321
|
+
- Visit upstream [agent-vision-toolkit](https://github.com/Anionex/agent-vision-toolkit) for the general toolkit, cross-agent integrations, and visual-task playbooks.
|
|
453
322
|
|
|
454
|
-
|
|
455
|
-
|
|
456
|
-
|
|
457
|
-
- Use [GitHub Issues](https://github.com/Anionex/dsh-vision-toolkit/issues) for reproducible bugs, focused feature requests, and usage questions; use [SUPPORT.md](SUPPORT.md) to choose the right channel.
|
|
458
|
-
- Report vulnerabilities privately through the process in [SECURITY.md](SECURITY.md), never in a public issue.
|
|
459
|
-
- Follow releases and compatibility notes in [CHANGELOG.md](CHANGELOG.md).
|
|
460
|
-
- Optional sponsorship is described transparently in [FUNDING.md](FUNDING.md); support does not purchase roadmap priority or private support.
|
|
461
|
-
- Use the upstream [project website](https://agent-vision.anionex.me) and [repository](https://github.com/Anionex/agent-vision-toolkit) for the general toolkit, cross-harness integrations, visual-task playbooks, and reference runs.
|
|
462
|
-
- Star, share, contribute to, or sponsor `agent-vision-toolkit` if its algorithms or methods save time; DSH-specific bugs and integration requests belong in this repository.
|
|
463
|
-
|
|
464
|
-
[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) was created by [Anionex](https://anionex.me/). This repository maintains its native DeepSeek Harness integration: DSH owns lifecycle, security, structured schemas, Credentials, Artifacts, and Web presentation, while the upstream project remains the home of the visual algorithms and reusable playbooks.
|
|
323
|
+
<p align="center">
|
|
324
|
+
<img src="assets/community-group-qr.png" alt="QR code for the agent-vision-toolkit community group" width="240" />
|
|
325
|
+
</p>
|
|
465
326
|
|
|
466
|
-
|
|
327
|
+
[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) was created by [Anionex](https://anionex.me/). This repository maintains its native DeepSeek Harness integration.
|
|
467
328
|
|
|
468
329
|
## License
|
|
469
330
|
|
|
470
|
-
The plugin is MIT
|
|
331
|
+
The plugin is available under the [MIT License](LICENSE). The packaged upstream snapshot retains its original MIT license in [`vendor/agent-vision-toolkit/LICENSE`](vendor/agent-vision-toolkit/LICENSE).
|