@anionex/dsh-vision-toolkit 0.1.13 → 0.1.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.i18n.yaml CHANGED
@@ -2,5 +2,5 @@
2
2
  # side as of the last confirmed-consistent state. Both languages carry equal authority;
3
3
  # after editing either side, bring the other along and re-record with:
4
4
  # pnpm run verify-translation-pairing --write dsh-vision-toolkit/README.md
5
- README.md: 5b09b23b96b986c9fdbbfccf7cd10b8728431d61
6
- README.zh.md: d3b383ebb44beb44c5513b7a03e31e7037bf98a9
5
+ README.md: 7beef0b9563811b385f5e49bc75ecaa8a21bc30f
6
+ README.zh.md: 93af175e51981174987f63c7c2e1f037f133293c
package/README.md CHANGED
@@ -7,21 +7,42 @@
7
7
  # DSH Vision Toolkit
8
8
 
9
9
  [![Recommended by dshfind](https://img.shields.io/badge/recommended%20by-dshfind-FFD700?style=flat-square)](https://dshfind.com/en/plugins/Anionex/dsh-vision-toolkit)
10
+ [![dshfind score: 94 — highest-rated plugin](https://img.shields.io/badge/dshfind%20score-94%20%7C%20highest--rated%20plugin-5B4CF0?style=flat-square)](https://dshfind.com/en/plugins/Anionex/dsh-vision-toolkit)
10
11
  [![npm](https://img.shields.io/npm/v/@anionex/dsh-vision-toolkit?style=flat-square&color=5B4CF0)](https://www.npmjs.com/package/@anionex/dsh-vision-toolkit)
11
12
  [![MIT](https://img.shields.io/badge/license-MIT-0B7285?style=flat-square)](LICENSE)
12
13
  [![DSH](https://img.shields.io/badge/DSH-Web%20%2B%20Headless-5B4CF0?style=flat-square)](cordis.patch.yml)
13
14
 
14
- **Give text-only DSH agents eyes: paste an image, ask a question, locate exact elements, extract assets, and verify UI restoration with measurable results.**
15
+ **A more powerful vision toolkit—give text-only models in DeepSeek Harness eyes: image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks in one toolkit and Skill.**
16
+
17
+ 🚀 Paste an image and ask directly | Install with one command | Built-in free quota | Broad use cases
15
18
 
16
19
  🌐 **English** | [中文](README.zh.md)
17
20
 
18
21
  </div>
19
22
 
20
- When you run DeepSeek or another text-only model in DeepSeek Harness (DSH), familiar problems appear quickly: the model cannot see a screenshot, generic image descriptions miss the point, buttons have no usable coordinates, and a rebuilt page may look “close enough” without any way to measure the remaining difference.
23
+ If you use DeepSeek or another text-only model in DeepSeek Harness (DSH), you may have run into the same problems: the model cannot see a screenshot, generic descriptions miss the point, buttons have no usable coordinates, and a rebuilt page can look “close enough” without a way to measure the remaining difference.
24
+
25
+ 🏆 This project is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem: it was initiated before internal beta and built during the beta with reference to [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit).
26
+
27
+ > **Original work:** This vision toolkit and the `vision-tools` Skill were personally created and continuously refined by the author through long-term real-world use and repeated iteration.
28
+
29
+ ## Highlights
30
+
31
+ - **Paste and use it immediately.** Paste an image in DSH Web and the text-only route switches to its `(Vision Toolkit)` variant without manual path copying or model changes.
32
+ - **A seamless image workflow.** Native thumbnails, session history, and workspace paths stay intact; Web can preview artifacts and Headless can continue using the same structured results.
33
+ - **One command to install.** The built-in free Groq Qwen3.6 vision service is ready after installation, with no API key required.
34
+ - **Built-in free quota.** The shared service includes 100 requests per client per day, 3,000 requests globally per day, and a 60-request burst per 60 seconds, with readable errors when a limit is reached.
35
+ - **Vision guided by intent.** The agent extracts evidence for the task at hand, such as “Where is the error?” or “Where is the button?”, instead of returning a generic caption.
36
+ - **A complete screenshot-to-verification loop.** Reference images, HTML screenshots, difference regions, and pixel comparison work together for UI restoration.
37
+
38
+ [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) gives an agent more than image captions: it can read, locate, crop, trace, rebuild, and verify visual work. DSH Vision Toolkit is its native DeepSeek Harness integration, bringing that workflow into Web and Headless Profiles.
21
39
 
22
- DSH Vision Toolkit packages [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) as a native DSH plugin. It helps an agent do more than describe an image: the agent can read, locate, crop, trace, rebuild, and verify visual work around the task at hand.
40
+ This project has two layers:
23
41
 
24
- > **Install and use it immediately.** The default setup includes a free Gemma 4 vision service and requires no API key. Cropping, pixel diffing, color analysis, foreground extraction, SVG tracing, and HTML screenshots run locally without spending vision API requests.
42
+ 1. **Visual tools and a Skill:** the agent learns when to inspect, ground, OCR, crop, trace, or compare pixels.
43
+ 2. **Native DSH integration:** those capabilities live inside Profiles, sessions, Settings, Artifacts, and the Web UI, with a free Groq Qwen3.6 vision service ready after installation.
44
+
45
+ > **Install and use it immediately.** The default setup includes a free Qwen3.6 vision service and requires no API key. Cropping, pixel diffing, color analysis, foreground extraction, SVG tracing, and HTML screenshots run locally without spending vision API requests.
25
46
 
26
47
  ```sh
27
48
  dsh plugin --profile web add @anionex/dsh-vision-toolkit
@@ -32,10 +53,10 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
32
53
  <details>
33
54
  <summary><strong>Table of contents</strong></summary>
34
55
 
56
+ - [Highlights](#highlights)
35
57
  - [Recent updates](#recent-updates)
36
- - [Problems it solves](#problems-it-solves)
58
+ - [Who it is for](#who-it-is-for)
37
59
  - [See it in action](#see-it-in-action)
38
- - [Highlights](#highlights)
39
60
  - [Quick start: three steps](#quick-start-three-steps)
40
61
  - [Common workflows](#common-workflows)
41
62
  - [Toolbox](#toolbox)
@@ -48,12 +69,12 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
48
69
  ## Recent updates
49
70
 
50
71
  - **2026-08-16 · Windows Python:** Added Microsoft Store Python support, fixing first-time isolated-runtime setup failures for affected Windows users.
51
- - **2026-08-16 · Better free vision:** Switched the default model to Gemma 4, improving the no-key image-understanding path.
72
+ - **2026-08-16 · Better free vision:** Switched the built-in no-key service to Groq Qwen3.6, improving image understanding without adding setup steps.
52
73
  - **2026-08-16 · Image paste:** Text-only routes now switch to a `(Vision Toolkit)` variant and keep a workspace path, fixing blocked pastes and images that could not be reused later.
53
- - **2026-08-16 · Higher free quotas:** Raised per-client, global, and burst limits to `100/day`, `400/day`, and `20/minute`, reducing avoidable rate-limit failures while the shared capacity is lightly used.
74
+ - **2026-08-16 · Higher free quotas:** Raised the shared service ceiling to `3,000/day` and `60/minute` to make better use of the three-account Groq pool while keeping the per-client limit at `100/day`.
54
75
  - **2026-08-16 · Real model test:** Added a full image-request test in Settings, fixing the false confidence caused by a successful `/models` request to a model that still cannot process images.
55
76
 
56
- ## Problems it solves
77
+ ## Who it is for
57
78
 
58
79
  | The problem | What Vision Toolkit delivers |
59
80
  |---|---|
@@ -94,22 +115,13 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
94
115
 
95
116
  ### Turn “looks close” into a verifiable result
96
117
 
97
- The repository includes a reproducible UI-restoration example. The first implementation differs from the reference by **6.04%**. After the highlighted regions are corrected, the final `1200 × 720` render reaches **0% pixel difference**.
118
+ The repository includes a reproducible UI-restoration example: the agent renders the reference and implementation, then uses difference regions, a heatmap, and a JSON report to guide the next correction.
98
119
 
99
120
  <p>
100
121
  <img src="examples/ui-restoration/assets/initial.png" width="49%" alt="Initial UI implementation with measurable layout and styling differences" />
101
- <img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="Final UI implementation after visual diagnosis, reaching zero pixel difference" />
122
+ <img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="UI implementation after visual diagnosis and pixel comparison" />
102
123
  </p>
103
124
 
104
- ## Highlights
105
-
106
- - **Free by default.** New installations use the built-in Gemma 4 service without requiring another account or API key.
107
- - **Focused on the current task.** The agent sends the reason it needs to inspect the image, so the result emphasizes useful evidence instead of producing a generic caption.
108
- - **Outputs you can keep working with.** Coordinates, OCR, transparent PNGs, SVGs, screenshots, heatmaps, and JSON can feed directly into the next step.
109
- - **Built for UI and screenshot work.** Reference analysis, element location, asset extraction, HTML rendering, and pixel comparison form one continuous workflow.
110
- - **Local where possible.** Crop, trace, pixel diff, color, foreground, and HTML screenshot operations do not need a remote vision model.
111
- - **The same capabilities in Web and Headless.** Web users can preview and download artifacts; Headless runs still receive replayable structured results and workspace paths.
112
-
113
125
  ## Quick start: three steps
114
126
 
115
127
  ### 1. Install
@@ -206,7 +218,7 @@ The default setup uses:
206
218
 
207
219
  ```text
208
220
  Base URL: https://vision.anionex.me/v1
209
- Model: gemma-4-26b-a4b-it
221
+ Model: qwen/qwen3.6-27b
210
222
  API Key: no user configuration required
211
223
  ```
212
224
 
@@ -215,8 +227,8 @@ This is a shared zero-configuration entry point, not an unlimited private endpoi
215
227
  | Limit | Current value |
216
228
  |---|---:|
217
229
  | Per client | 100 requests per UTC day |
218
- | Whole service | 400 requests per UTC day |
219
- | Burst | 20 requests per 60 seconds |
230
+ | Whole service | 3,000 requests per UTC day |
231
+ | Burst | 60 requests per 60 seconds |
220
232
  | Image size | 4 MiB per image |
221
233
  | Decoded pixels | 20,000,000 per image |
222
234
  | Output | 512 tokens per request |
package/README.zh.md CHANGED
@@ -7,21 +7,47 @@
7
7
  # DSH Vision Toolkit
8
8
 
9
9
  [![由 dshfind 推荐](https://img.shields.io/badge/%E7%94%B1%20dshfind-%E6%8E%A8%E8%8D%90-FFD700?style=flat-square)](https://dshfind.com/zh/plugins/Anionex/dsh-vision-toolkit)
10
+ [![dshfind 评分:94——最高分插件](https://img.shields.io/badge/dshfind%20%E8%AF%84%E5%88%86-94%20%7C%20%E6%9C%80%E9%AB%98%E5%88%86%E6%8F%92%E4%BB%B6-5B4CF0?style=flat-square)](https://dshfind.com/zh/plugins/Anionex/dsh-vision-toolkit)
10
11
  [![npm](https://img.shields.io/npm/v/@anionex/dsh-vision-toolkit?style=flat-square&color=5B4CF0)](https://www.npmjs.com/package/@anionex/dsh-vision-toolkit)
11
12
  [![MIT](https://img.shields.io/badge/license-MIT-0B7285?style=flat-square)](LICENSE)
12
13
  [![DSH](https://img.shields.io/badge/DSH-Web%20%2B%20Headless-5B4CF0?style=flat-square)](cordis.patch.yml)
13
14
 
14
- **给 DSH 里的纯文本 Agent 装上眼睛:粘贴图片就能问,找到元素就能继续操作,还能把 UI 还原做到有数据可验。**
15
+ **更强大的视觉工具箱——给 DeepSeek Harness 里的纯文本模型装上眼睛:图片问答、长图 OCR、前端 UI 还原、GUI 视觉任务,一套视觉工具箱和一个 Skill。**
16
+
17
+ 🚀 粘贴图片,直接提问 | 一行命令安装即用 | 内置免费额度 | 场景丰富
15
18
 
16
19
  🌐 [English](README.md) | **中文**
17
20
 
18
21
  </div>
19
22
 
20
- 如果你在 DeepSeek Harness(DSH)里使用 DeepSeek 等纯文本模型,可能已经遇到过这些问题:模型看不到截图、只能得到一段泛泛的图片描述、找不到按钮的准确位置,或者还原出来的页面“看起来差不多”,却不知道到底差了多少。
23
+ 如果你在 dsh 中使用 DeepSeek 等纯文本模型,遇到了下面问题中的一个或者多个,那么这个插件适合你:
24
+ 1. 粘贴图片被拒绝,不能发图片给模型,还要手动切换模型。
25
+ 2. 模型看不到图片内容,不能做和图片有关的任务。
26
+ 3. 已有方案只能得到图片笼统描述,完成不了高难度视觉相关任务,例如ui还原,长截图分析等。
27
+ 4. 不能安装即用,直接体验,还要自己配置api key。
28
+
29
+ 🏆 本项目为deepseek harness生态首个综合性视觉工具插件:内测前已立项,并在内测期间参考本人的[`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit)做出
30
+
31
+ > **原创声明:** 这套视觉工具和 `vision-tools` Skill 由作者个人原创并持续打磨,相关工具、方法和工作流来自长期的真实使用与反复迭代。
32
+
33
+ ## 亮点
34
+
35
+ - **粘贴即可使用。** 在 DSH Web 里粘贴图片,文本模型会自动切换到看图模式变体,不需要手动复制路径或更换模型。
36
+ - **无缝体验。** 图片保留原生缩略图、会话记录和工作区路径;Web 可以预览产物,Headless 也能继续使用同一份结构化结果。
37
+ - **一行命令安装即用。** 安装插件后默认使用内置免费 Groq Qwen3.6 视觉服务,不需要申请 API Key。
38
+ - **内置免费额度。** 共享服务提供每客户端每天 100 次、全局每天 3,000 次和每 60 秒 60 次突发请求额度,并在触发限制时返回明确原因。
39
+ - **带着意图去看图。** Agent 不只生成通用描述,而是围绕“报错在哪里”“按钮在哪”等当前任务提取证据。
40
+ - **从截图到可验证结果。** 参考图、HTML 截图、差异定位和像素对比组成一条完整 UI 还原闭环。
41
+
21
42
 
22
- DSH Vision Toolkit 把 [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) 变成一个原生 DSH 插件。安装后,Agent 不只会“看图”,还会围绕当前任务读取、定位、裁剪、描摹、还原和验证图片。
43
+ [`agent-vision-toolkit`](https://github.com/Anionex/agent-vision-toolkit) 的视觉能力不只停留在图片描述:Agent 可以读取、定位、裁剪、描摹、还原和验证视觉内容。DSH Vision Toolkit 是这套工具箱面向 DeepSeek Harness 的原生接入,让它进入 Web 和 Headless Profile。
23
44
 
24
- > **安装即可使用。** 默认接入内置免费 Gemma 4 视觉服务,不需要申请 API Key;裁图、像素对比、颜色分析、前景提取、SVG 描摹和网页截图等本地工具也不消耗视觉 API 请求。
45
+ 本项目提供两层能力:
46
+
47
+ 1. **视觉工具和 Skill**:让 Agent 知道什么时候该看图、定位、OCR、裁剪、描摹或做像素对比。
48
+ 2. **DSH 原生接入**:把这些能力放进 Profile、会话、Settings、Artifacts 和 Web 界面,并提供安装即可使用的免费 Groq Qwen3.6 视觉服务。
49
+
50
+ > **安装即可使用。** 默认接入内置免费 Groq Qwen3.6 视觉服务,不需要申请 API Key;
25
51
 
26
52
  ```sh
27
53
  dsh plugin --profile web add @anionex/dsh-vision-toolkit
@@ -32,10 +58,10 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
32
58
  <details>
33
59
  <summary><strong>目录</strong></summary>
34
60
 
61
+ - [亮点](#亮点)
35
62
  - [最近更新](#最近更新)
36
- - [它解决什么问题](#它解决什么问题)
63
+ - [适合谁用](#适合谁用)
37
64
  - [实际效果](#实际效果)
38
- - [亮点](#亮点)
39
65
  - [快速开始:三步完成](#快速开始三步完成)
40
66
  - [常见任务](#常见任务)
41
67
  - [工具一览](#工具一览)
@@ -48,12 +74,12 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
48
74
  ## 最近更新
49
75
 
50
76
  - **2026-08-16 · Windows Python:** 支持 Microsoft Store Python,解决部分 Windows 用户首次创建隔离环境失败的问题。
51
- - **2026-08-16 · 免费视觉升级:** 默认模型切换到 Gemma 4,解决免 Key 方案看图效果不足的问题。
77
+ - **2026-08-16 · 免费视觉升级:** 默认模型切换到 Groq Qwen3.6,解决免 Key 方案看图效果不足的问题。
52
78
  - **2026-08-16 · 图片粘贴:** 文本模型自动切换到 `(Vision Toolkit)` 变体并保留工作区路径,解决粘贴图片被拦截或后续无法复用的问题。
53
- - **2026-08-16 · 免费额度:** 单客户端、全局和突发额度分别提高到 `100/日`、`400/日` 和 `20/分钟`,解决早期用户容易撞限而共享额度闲置的问题。
79
+ - **2026-08-16 · 免费额度:** 全局和突发额度提高到 `3,000/日` 和 `60/分钟`,更充分利用三个 Groq 账号,同时保留单客户端 `100/日` 限制。
54
80
  - **2026-08-16 · 真实模型测试:** Settings 新增完整图片请求测试,解决 `/models` 可访问却不能证明模型真的会看图的问题。
55
81
 
56
- ## 它解决什么问题
82
+ ## 适合谁用
57
83
 
58
84
  | 你遇到的问题 | Vision Toolkit 给出的结果 |
59
85
  |---|---|
@@ -94,22 +120,13 @@ dsh plugin --profile web add @anionex/dsh-vision-toolkit
94
120
 
95
121
  ### 让“差不多”变成“可验证”
96
122
 
97
- 仓库内置了一个可复现的 UI 还原示例:初版与参考图的差异为 **6.04%**,经过定位和修正后,在 `1200 × 720` 下达到 **0% 像素差异**。
123
+ 仓库内置了一个可复现的 UI 还原示例:Agent 会先渲染参考图和实现,再用差异区域、热力图和 JSON 报告指导下一轮修正。
98
124
 
99
125
  <p>
100
126
  <img src="examples/ui-restoration/assets/initial.png" width="49%" alt="像素对比前仍有布局和样式偏差的初版 UI" />
101
- <img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="经过视觉定位和像素对比后达到零差异的最终 UI" />
127
+ <img src="examples/ui-restoration/assets/implementation.png" width="49%" alt="经过视觉定位和像素对比后的 UI 实现" />
102
128
  </p>
103
129
 
104
- ## 亮点
105
-
106
- - **安装后就能免费用。** 新用户默认使用内置 Gemma 4 服务,不需要注册新的模型平台,也不需要先填写 Key。
107
- - **不只描述图片,而是解决当前问题。** Agent 会把当前任务作为视觉关注点,优先返回这一轮真正要用到的内容。
108
- - **返回可以继续工作的结果。** 坐标、OCR、透明 PNG、SVG、截图、热力图和 JSON 都能直接交给下一步。
109
- - **特别适合 UI 和截图工程。** 从参考图、元素定位、素材提取到 HTML 截图和像素对比,形成完整闭环。
110
- - **能本地做的就本地做。** 裁剪、描摹、像素对比、颜色、前景和 HTML 截图不需要上传到视觉模型。
111
- - **Web 与 Headless 使用同一套能力。** Web 中可以预览和下载产物,Headless 中仍会得到可重放的结构化结果和文件路径。
112
-
113
130
  ## 快速开始:三步完成
114
131
 
115
132
  ### 1. 安装
@@ -206,7 +223,7 @@ flowchart LR
206
223
 
207
224
  ```text
208
225
  Base URL: https://vision.anionex.me/v1
209
- Model: gemma-4-26b-a4b-it
226
+ Model: qwen/qwen3.6-27b
210
227
  API Key: 不需要用户配置
211
228
  ```
212
229
 
@@ -215,8 +232,8 @@ API Key: 不需要用户配置
215
232
  | 限制 | 当前值 |
216
233
  |---|---:|
217
234
  | 单客户端 | 每个 UTC 日 100 次 |
218
- | 全局服务 | 每个 UTC 日 400 次 |
219
- | 突发请求 | 60 秒内 20 次 |
235
+ | 全局服务 | 每个 UTC 日 3,000 次 |
236
+ | 突发请求 | 60 秒内 60 次 |
220
237
  | 单张图片大小 | 4 MiB |
221
238
  | 单张图片像素 | 20,000,000 |
222
239
  | 单次输出 | 512 tokens |
package/lib/client.js CHANGED
@@ -21,7 +21,7 @@ const DEFAULT_USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKi
21
21
  // Keep these browser defaults aligned with src/defaults.ts without importing server-side config.
22
22
  const BUILT_IN_FREE_VISION_BASE_URL = 'https://vision.anionex.me/v1';
23
23
  const BUILT_IN_FREE_VISION_CREDENTIAL = 'ANIONEX_FREE_VISION';
24
- const BUILT_IN_FREE_VISION_MODEL = 'gemma-4-26b-a4b-it';
24
+ const BUILT_IN_FREE_VISION_MODEL = 'qwen/qwen3.6-27b';
25
25
  const en = {
26
26
  nav: 'Vision',
27
27
  settingsTitle: 'Vision Toolkit',