@liustack/modlens 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +77 -0
- package/README.zh-CN.md +69 -0
- package/dist/main.js +275 -0
- package/package.json +54 -0
- package/skills/modlens/SKILL.md +75 -0
- package/skills/modlens/references/output-schema.md +57 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Leon Liu (liustack)
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,77 @@
|
|
|
1
|
+
# ModLens
|
|
2
|
+
|
|
3
|
+
A CLI toolkit for AI agents that converts image sources (local path or remote URL) into structured text evidence, bridging the vision gap for text-only LLM workflows.
|
|
4
|
+
|
|
5
|
+
[中文说明](README.zh-CN.md)
|
|
6
|
+
|
|
7
|
+
## Features
|
|
8
|
+
|
|
9
|
+
- Built for non-vision LLM setups (text-only models + external vision bridge)
|
|
10
|
+
- Supports local image paths and remote image URLs
|
|
11
|
+
- Pluggable vision backend — ships with Gemini CLI; more engines (PaddleOCR, DeepSeek, etc.) planned
|
|
12
|
+
- Outputs machine-consumable JSON (OCR + layout + semantics + visual clues)
|
|
13
|
+
- Designed to be called from Agent Skills (Claude Code, Codex, Cursor, etc.)
|
|
14
|
+
|
|
15
|
+
## Install
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
npm install -g @liustack/modlens
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
The default backend requires Gemini CLI to be installed and authenticated:
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
npm install -g @google/gemini-cli
|
|
25
|
+
gemini
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Or run with `npx`:
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
npx @liustack/modlens [options]
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
## Usage
|
|
35
|
+
|
|
36
|
+
```bash
|
|
37
|
+
# Print JSON result to stdout
|
|
38
|
+
modlens -i screenshot.png
|
|
39
|
+
|
|
40
|
+
# Save to file
|
|
41
|
+
modlens -i screenshot.png -o lens.json
|
|
42
|
+
|
|
43
|
+
# Specify model + extra prompt constraints
|
|
44
|
+
modlens -i screenshot.png -m gemini-2.5-flash --prompt "Focus on table structure"
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Options
|
|
48
|
+
|
|
49
|
+
| Flag | Description |
|
|
50
|
+
|------|-------------|
|
|
51
|
+
| `-i, --input <path>` | Input image path (required) |
|
|
52
|
+
| `-o, --output <path>` | Write result JSON to a file |
|
|
53
|
+
| `-m, --model <name>` | Vision model name (backend-specific) |
|
|
54
|
+
| `--prompt <text>` | Extra extraction constraints |
|
|
55
|
+
| `--timeout <ms>` | Timeout in milliseconds (default: `180000`) |
|
|
56
|
+
| `--gemini-bin <path>` | Gemini CLI binary path (default: `gemini`) |
|
|
57
|
+
|
|
58
|
+
## Vision Backends
|
|
59
|
+
|
|
60
|
+
ModLens uses a pluggable architecture for vision recognition. The current v1 ships with **Gemini CLI** as the default backend. Future versions will support additional engines such as PaddleOCR, DeepSeek OCR, and other multimodal/vision-capable models.
|
|
61
|
+
|
|
62
|
+
## Agent Skill
|
|
63
|
+
|
|
64
|
+
- [modlens/SKILL.md](skills/modlens/SKILL.md)
|
|
65
|
+
|
|
66
|
+
## Notes
|
|
67
|
+
|
|
68
|
+
- `modlens` focuses on visual parsing only.
|
|
69
|
+
- `modsearch` and `modfetch` belong to separate projects and are intentionally out of scope.
|
|
70
|
+
|
|
71
|
+
## Disclaimer
|
|
72
|
+
|
|
73
|
+
This project is for **personal learning and experimentation only**. It is not intended for commercial use.
|
|
74
|
+
|
|
75
|
+
## License
|
|
76
|
+
|
|
77
|
+
MIT
|
package/README.zh-CN.md
ADDED
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# ModLens
|
|
2
|
+
|
|
3
|
+
面向 AI Agent 的视觉外挂 CLI,用于把图片来源(本地路径或远程 URL)转成结构化文本证据,补齐纯文本 LLM 的“看图”能力。
|
|
4
|
+
|
|
5
|
+
## 特性
|
|
6
|
+
|
|
7
|
+
- 面向“无视觉能力模型”场景(文本模型 + 外挂视觉)
|
|
8
|
+
- 支持本地图片路径与远程图片 URL
|
|
9
|
+
- 基于 Gemini CLI 非交互调用(`gemini -p`)
|
|
10
|
+
- 输出结构化 JSON(OCR + 布局 + 语义 + 视觉线索)
|
|
11
|
+
- 适合作为 Agent Skill 工具被 Claude Code / Codex 等调用
|
|
12
|
+
|
|
13
|
+
## 安装
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
npm install -g @liustack/modlens
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
需要先安装并认证 Gemini CLI:
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
npm install -g @google/gemini-cli
|
|
23
|
+
gemini
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
或直接用 `npx`:
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
npx @liustack/modlens [options]
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## 用法
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
# 标准输出 JSON
|
|
36
|
+
modlens -i screenshot.png
|
|
37
|
+
|
|
38
|
+
# 落盘到文件
|
|
39
|
+
modlens -i screenshot.png -o lens.json
|
|
40
|
+
|
|
41
|
+
# 指定模型和额外解析要求
|
|
42
|
+
modlens -i screenshot.png -m gemini-2.5-flash --prompt "重点提取表格结构"
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
## 参数
|
|
46
|
+
|
|
47
|
+
- `-i, --input <path>` 输入图片路径(必填)
|
|
48
|
+
- `-o, --output <path>` 可选输出 JSON 路径
|
|
49
|
+
- `-m, --model <name>` Gemini 模型名
|
|
50
|
+
- `--prompt <text>` 额外解析约束
|
|
51
|
+
- `--timeout <ms>` 超时毫秒(默认 `180000`)
|
|
52
|
+
- `--gemini-bin <path>` Gemini 可执行路径(默认 `gemini`)
|
|
53
|
+
|
|
54
|
+
## Agent Skill
|
|
55
|
+
|
|
56
|
+
- [modlens/SKILL.md](skills/modlens/SKILL.md)
|
|
57
|
+
|
|
58
|
+
## 说明
|
|
59
|
+
|
|
60
|
+
- `modlens` 只做视觉解析。
|
|
61
|
+
- `modsearch` / `modfetch` 是其他独立项目,不在本仓库实现。
|
|
62
|
+
|
|
63
|
+
## 免责声明
|
|
64
|
+
|
|
65
|
+
本项目仅供**个人学习与实验**使用,请勿用于商业用途。
|
|
66
|
+
|
|
67
|
+
## License
|
|
68
|
+
|
|
69
|
+
MIT
|
package/dist/main.js
ADDED
|
@@ -0,0 +1,275 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
import { Command } from "commander";
|
|
3
|
+
import * as fs from "fs";
|
|
4
|
+
import * as path from "path";
|
|
5
|
+
import { spawn } from "child_process";
|
|
6
|
+
function buildVisionPrompt(imageAbsolutePath, extraPrompt) {
|
|
7
|
+
const basePrompt = `Analyze this image: ${imageAbsolutePath}
|
|
8
|
+
|
|
9
|
+
You are an image parsing engine.
|
|
10
|
+
Analyze the provided image.
|
|
11
|
+
|
|
12
|
+
Goal: Convert all image information into structured results consumable by a text-only LLM.
|
|
13
|
+
|
|
14
|
+
Strict requirements:
|
|
15
|
+
1. Output JSON only. No Markdown, no explanatory text.
|
|
16
|
+
2. Cover all visible text, semantics, structure, layout, and visual clues as thoroughly as possible.
|
|
17
|
+
3. If any information is uncertain, note it in the uncertainty field.
|
|
18
|
+
|
|
19
|
+
Output JSON structure:
|
|
20
|
+
{
|
|
21
|
+
"summary": "",
|
|
22
|
+
"ocr": {
|
|
23
|
+
"full_text": "",
|
|
24
|
+
"lines": [
|
|
25
|
+
{
|
|
26
|
+
"text": "",
|
|
27
|
+
"language": "",
|
|
28
|
+
"confidence": 0
|
|
29
|
+
}
|
|
30
|
+
]
|
|
31
|
+
},
|
|
32
|
+
"layout": {
|
|
33
|
+
"regions": [
|
|
34
|
+
{
|
|
35
|
+
"id": "",
|
|
36
|
+
"type": "title|subtitle|paragraph|list|table|chart|form|image|icon|other",
|
|
37
|
+
"bbox": { "x": 0, "y": 0, "w": 0, "h": 0 },
|
|
38
|
+
"reading_order": 1,
|
|
39
|
+
"text": ""
|
|
40
|
+
}
|
|
41
|
+
]
|
|
42
|
+
},
|
|
43
|
+
"semantics": {
|
|
44
|
+
"scene": "",
|
|
45
|
+
"intent": "",
|
|
46
|
+
"entities": [
|
|
47
|
+
{
|
|
48
|
+
"name": "",
|
|
49
|
+
"type": "",
|
|
50
|
+
"evidence": ""
|
|
51
|
+
}
|
|
52
|
+
],
|
|
53
|
+
"relations": [
|
|
54
|
+
{
|
|
55
|
+
"subject": "",
|
|
56
|
+
"predicate": "",
|
|
57
|
+
"object": ""
|
|
58
|
+
}
|
|
59
|
+
]
|
|
60
|
+
},
|
|
61
|
+
"visual": {
|
|
62
|
+
"dominant_colors": [""],
|
|
63
|
+
"style": "",
|
|
64
|
+
"notes": [""]
|
|
65
|
+
},
|
|
66
|
+
"uncertainty": [""]
|
|
67
|
+
}`;
|
|
68
|
+
if (!extraPrompt || !extraPrompt.trim()) {
|
|
69
|
+
return basePrompt;
|
|
70
|
+
}
|
|
71
|
+
return `${basePrompt}
|
|
72
|
+
|
|
73
|
+
Additional requirements:
|
|
74
|
+
${extraPrompt.trim()}`;
|
|
75
|
+
}
|
|
76
|
+
const DEFAULT_TIMEOUT_MS = 18e4;
|
|
77
|
+
function buildGeminiInvocation(options) {
|
|
78
|
+
const isRemote = isRemoteSource(options.imagePath);
|
|
79
|
+
const imageSource = isRemote ? options.imagePath.trim() : path.resolve(options.imagePath);
|
|
80
|
+
const imageDir = path.dirname(imageSource);
|
|
81
|
+
const prompt = buildVisionPrompt(imageSource, options.extraPrompt);
|
|
82
|
+
const args = ["-p", prompt, "--output-format", "json"];
|
|
83
|
+
if (options.model) {
|
|
84
|
+
args.push("-m", options.model);
|
|
85
|
+
}
|
|
86
|
+
return {
|
|
87
|
+
command: options.geminiBin || "gemini",
|
|
88
|
+
args,
|
|
89
|
+
cwd: options.workspaceDir || (isRemote ? process.cwd() : imageDir)
|
|
90
|
+
};
|
|
91
|
+
}
|
|
92
|
+
function parseGeminiCliJsonOutput(stdout) {
|
|
93
|
+
let parsed;
|
|
94
|
+
try {
|
|
95
|
+
parsed = JSON.parse(stdout.trim());
|
|
96
|
+
} catch (error) {
|
|
97
|
+
throw new Error(`Failed to parse Gemini JSON output: ${error.message}`);
|
|
98
|
+
}
|
|
99
|
+
if (!parsed || typeof parsed !== "object" || typeof parsed.response !== "string") {
|
|
100
|
+
throw new Error("Gemini JSON output is missing a string `response` field.");
|
|
101
|
+
}
|
|
102
|
+
return parsed;
|
|
103
|
+
}
|
|
104
|
+
function extractStructuredResponse(text) {
|
|
105
|
+
const rawText = text.trim();
|
|
106
|
+
const direct = tryParseJson(rawText);
|
|
107
|
+
if (direct !== null) {
|
|
108
|
+
return {
|
|
109
|
+
structured: direct,
|
|
110
|
+
rawText
|
|
111
|
+
};
|
|
112
|
+
}
|
|
113
|
+
const fencedMatch = /```(?:json)?\s*([\s\S]*?)```/i.exec(rawText);
|
|
114
|
+
if (fencedMatch) {
|
|
115
|
+
const parsedFenced = tryParseJson(fencedMatch[1].trim());
|
|
116
|
+
if (parsedFenced !== null) {
|
|
117
|
+
return {
|
|
118
|
+
structured: parsedFenced,
|
|
119
|
+
rawText
|
|
120
|
+
};
|
|
121
|
+
}
|
|
122
|
+
}
|
|
123
|
+
const firstBrace = rawText.indexOf("{");
|
|
124
|
+
const lastBrace = rawText.lastIndexOf("}");
|
|
125
|
+
if (firstBrace >= 0 && lastBrace > firstBrace) {
|
|
126
|
+
const possibleJson = rawText.slice(firstBrace, lastBrace + 1);
|
|
127
|
+
const parsedObject = tryParseJson(possibleJson);
|
|
128
|
+
if (parsedObject !== null) {
|
|
129
|
+
return {
|
|
130
|
+
structured: parsedObject,
|
|
131
|
+
rawText
|
|
132
|
+
};
|
|
133
|
+
}
|
|
134
|
+
}
|
|
135
|
+
return {
|
|
136
|
+
structured: null,
|
|
137
|
+
rawText
|
|
138
|
+
};
|
|
139
|
+
}
|
|
140
|
+
async function analyzeImage(options) {
|
|
141
|
+
const resolvedInput = resolveInput(options.input);
|
|
142
|
+
if (resolvedInput.kind === "local") {
|
|
143
|
+
validateInputFile(resolvedInput.source);
|
|
144
|
+
}
|
|
145
|
+
const invocation = buildGeminiInvocation({
|
|
146
|
+
imagePath: resolvedInput.source,
|
|
147
|
+
model: options.model,
|
|
148
|
+
geminiBin: options.geminiBin,
|
|
149
|
+
workspaceDir: options.workspaceDir,
|
|
150
|
+
extraPrompt: options.prompt
|
|
151
|
+
});
|
|
152
|
+
const commandResult = await runCommand(invocation, options.timeoutMs ?? DEFAULT_TIMEOUT_MS);
|
|
153
|
+
const geminiOutput = parseGeminiCliJsonOutput(commandResult.stdout);
|
|
154
|
+
const extracted = extractStructuredResponse(geminiOutput.response);
|
|
155
|
+
return {
|
|
156
|
+
image: resolvedInput.source,
|
|
157
|
+
structured: extracted.structured,
|
|
158
|
+
rawText: extracted.rawText,
|
|
159
|
+
meta: {
|
|
160
|
+
generatedAt: (/* @__PURE__ */ new Date()).toISOString(),
|
|
161
|
+
model: options.model ?? null,
|
|
162
|
+
geminiSessionId: geminiOutput.session_id ?? null,
|
|
163
|
+
geminiStats: geminiOutput.stats ?? null
|
|
164
|
+
}
|
|
165
|
+
};
|
|
166
|
+
}
|
|
167
|
+
function resolveInput(input) {
|
|
168
|
+
const trimmed = input.trim();
|
|
169
|
+
if (!trimmed) {
|
|
170
|
+
throw new Error("Input path is required.");
|
|
171
|
+
}
|
|
172
|
+
if (isRemoteSource(trimmed)) {
|
|
173
|
+
return { source: trimmed, kind: "remote" };
|
|
174
|
+
}
|
|
175
|
+
if (/^file:\/\//i.test(trimmed)) {
|
|
176
|
+
const localPath = decodeURI(trimmed.replace(/^file:\/\//i, ""));
|
|
177
|
+
return { source: path.resolve(localPath), kind: "local" };
|
|
178
|
+
}
|
|
179
|
+
return { source: path.resolve(trimmed), kind: "local" };
|
|
180
|
+
}
|
|
181
|
+
function isRemoteSource(value) {
|
|
182
|
+
return /^https?:\/\//i.test(value.trim());
|
|
183
|
+
}
|
|
184
|
+
function validateInputFile(filePath) {
|
|
185
|
+
if (!fs.existsSync(filePath)) {
|
|
186
|
+
throw new Error(`Input image not found: ${filePath}`);
|
|
187
|
+
}
|
|
188
|
+
const stat = fs.statSync(filePath);
|
|
189
|
+
if (!stat.isFile()) {
|
|
190
|
+
throw new Error(`Input is not a file: ${filePath}`);
|
|
191
|
+
}
|
|
192
|
+
}
|
|
193
|
+
function tryParseJson(text) {
|
|
194
|
+
try {
|
|
195
|
+
return JSON.parse(text);
|
|
196
|
+
} catch {
|
|
197
|
+
return null;
|
|
198
|
+
}
|
|
199
|
+
}
|
|
200
|
+
function runCommand(invocation, timeoutMs) {
|
|
201
|
+
return new Promise((resolve, reject) => {
|
|
202
|
+
const child = spawn(invocation.command, invocation.args, {
|
|
203
|
+
cwd: invocation.cwd,
|
|
204
|
+
stdio: ["ignore", "pipe", "pipe"]
|
|
205
|
+
});
|
|
206
|
+
let stdout = "";
|
|
207
|
+
let stderr = "";
|
|
208
|
+
let timedOut = false;
|
|
209
|
+
const timer = setTimeout(() => {
|
|
210
|
+
timedOut = true;
|
|
211
|
+
child.kill("SIGTERM");
|
|
212
|
+
}, timeoutMs);
|
|
213
|
+
child.stdout.on("data", (chunk) => {
|
|
214
|
+
stdout += chunk.toString();
|
|
215
|
+
});
|
|
216
|
+
child.stderr.on("data", (chunk) => {
|
|
217
|
+
stderr += chunk.toString();
|
|
218
|
+
});
|
|
219
|
+
child.on("error", (error) => {
|
|
220
|
+
clearTimeout(timer);
|
|
221
|
+
if (error.code === "ENOENT") {
|
|
222
|
+
reject(new Error(`Gemini CLI not found: ${invocation.command}`));
|
|
223
|
+
return;
|
|
224
|
+
}
|
|
225
|
+
reject(error);
|
|
226
|
+
});
|
|
227
|
+
child.on("close", (code) => {
|
|
228
|
+
clearTimeout(timer);
|
|
229
|
+
if (timedOut) {
|
|
230
|
+
reject(new Error(`Gemini CLI timed out after ${timeoutMs} ms.`));
|
|
231
|
+
return;
|
|
232
|
+
}
|
|
233
|
+
if (code !== 0) {
|
|
234
|
+
reject(
|
|
235
|
+
new Error(
|
|
236
|
+
`Gemini CLI failed with code ${code}.${stderr ? ` stderr: ${stderr.trim()}` : ""}`
|
|
237
|
+
)
|
|
238
|
+
);
|
|
239
|
+
return;
|
|
240
|
+
}
|
|
241
|
+
resolve({ stdout, stderr });
|
|
242
|
+
});
|
|
243
|
+
});
|
|
244
|
+
}
|
|
245
|
+
const program = new Command();
|
|
246
|
+
program.name("modlens").description("Image-to-text visual bridge for non-vision LLM workflows").version("0.1.0").requiredOption("-i, --input <path>", "Input image path").option("-o, --output <path>", "Write result JSON to a file").option("-m, --model <name>", "Gemini model name").option("--prompt <text>", "Extra prompt constraints for this image").option("--timeout <ms>", "Command timeout in milliseconds", "180000").option("--gemini-bin <path>", "Gemini CLI binary path", "gemini").action(async (options) => {
|
|
247
|
+
try {
|
|
248
|
+
const timeoutMs = Number.parseInt(options.timeout, 10);
|
|
249
|
+
if (!Number.isFinite(timeoutMs) || timeoutMs <= 0) {
|
|
250
|
+
throw new Error("Invalid --timeout. Use a positive integer in milliseconds.");
|
|
251
|
+
}
|
|
252
|
+
const result = await analyzeImage({
|
|
253
|
+
input: options.input,
|
|
254
|
+
model: options.model,
|
|
255
|
+
prompt: options.prompt,
|
|
256
|
+
timeoutMs,
|
|
257
|
+
geminiBin: options.geminiBin
|
|
258
|
+
});
|
|
259
|
+
const output = JSON.stringify(result, null, 2);
|
|
260
|
+
if (options.output) {
|
|
261
|
+
const outputPath = path.resolve(options.output);
|
|
262
|
+
fs.mkdirSync(path.dirname(outputPath), { recursive: true });
|
|
263
|
+
fs.writeFileSync(outputPath, output, "utf-8");
|
|
264
|
+
}
|
|
265
|
+
process.stdout.write(`${output}
|
|
266
|
+
`);
|
|
267
|
+
} catch (error) {
|
|
268
|
+
process.stderr.write(
|
|
269
|
+
`Error: ${error instanceof Error ? error.message : String(error)}
|
|
270
|
+
`
|
|
271
|
+
);
|
|
272
|
+
process.exit(1);
|
|
273
|
+
}
|
|
274
|
+
});
|
|
275
|
+
program.parse();
|
package/package.json
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "@liustack/modlens",
|
|
3
|
+
"version": "0.1.0",
|
|
4
|
+
"description": "CLI tool to provide visual understanding for non-vision LLMs",
|
|
5
|
+
"type": "module",
|
|
6
|
+
"bin": {
|
|
7
|
+
"modlens": "./dist/main.js"
|
|
8
|
+
},
|
|
9
|
+
"files": [
|
|
10
|
+
"dist",
|
|
11
|
+
"skills/modlens/SKILL.md",
|
|
12
|
+
"skills/modlens/references"
|
|
13
|
+
],
|
|
14
|
+
"keywords": [
|
|
15
|
+
"cli",
|
|
16
|
+
"vision",
|
|
17
|
+
"ocr",
|
|
18
|
+
"gemini",
|
|
19
|
+
"claude-code",
|
|
20
|
+
"skill",
|
|
21
|
+
"image-to-text",
|
|
22
|
+
"multimodal",
|
|
23
|
+
"modlens"
|
|
24
|
+
],
|
|
25
|
+
"author": "Leon Liu",
|
|
26
|
+
"license": "MIT",
|
|
27
|
+
"repository": {
|
|
28
|
+
"type": "git",
|
|
29
|
+
"url": "git+https://github.com/liustack/modlens.git"
|
|
30
|
+
},
|
|
31
|
+
"bugs": {
|
|
32
|
+
"url": "https://github.com/liustack/modlens/issues"
|
|
33
|
+
},
|
|
34
|
+
"homepage": "https://github.com/liustack/modlens#readme",
|
|
35
|
+
"engines": {
|
|
36
|
+
"node": ">=18"
|
|
37
|
+
},
|
|
38
|
+
"dependencies": {
|
|
39
|
+
"commander": "^13.1.0"
|
|
40
|
+
},
|
|
41
|
+
"devDependencies": {
|
|
42
|
+
"@types/node": "^22.19.7",
|
|
43
|
+
"typescript": "^5.9.3",
|
|
44
|
+
"vite": "^6.4.1",
|
|
45
|
+
"vitest": "^3.2.4"
|
|
46
|
+
},
|
|
47
|
+
"scripts": {
|
|
48
|
+
"dev": "vite build --watch",
|
|
49
|
+
"build": "vite build",
|
|
50
|
+
"typecheck": "tsc --noEmit",
|
|
51
|
+
"test": "vitest run",
|
|
52
|
+
"docs:list": "node scripts/docs-list.js"
|
|
53
|
+
}
|
|
54
|
+
}
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: modlens
|
|
3
|
+
description: "Bridge image understanding for non-vision LLM workflows. Use when user provides screenshots/photos/charts/doc images and the active model lacks multimodal vision. Call modlens to extract OCR text, semantics, structure, and layout as JSON evidence."
|
|
4
|
+
allowed-tools:
|
|
5
|
+
- Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# ModLens — Vision Bridge Skill
|
|
9
|
+
|
|
10
|
+
Use this skill when:
|
|
11
|
+
- User asks to analyze an image/screenshot/chart/document photo
|
|
12
|
+
- Current model cannot directly read images
|
|
13
|
+
- You need structured visual evidence before downstream reasoning
|
|
14
|
+
|
|
15
|
+
Do not use this skill for:
|
|
16
|
+
- Web search (`modsearch`)
|
|
17
|
+
- Web fetch (`modfetch`)
|
|
18
|
+
|
|
19
|
+
## Prerequisites
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
modlens --version
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
The default vision backend (Gemini CLI) must also be installed and authenticated:
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
gemini --version
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
If `gemini` is missing:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
npm install -g @google/gemini-cli
|
|
35
|
+
gemini
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
## Command
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
modlens -i <image-path>
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Optional:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
modlens -i <image-path> -o <output-json-path> -m <model-name> --prompt "<extra constraints>"
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
## Workflow
|
|
51
|
+
|
|
52
|
+
1. If user includes one or more images, run `modlens` for each image.
|
|
53
|
+
2. Parse returned JSON.
|
|
54
|
+
3. Feed `summary`, `ocr`, `layout`, and `semantics` back into your reasoning context.
|
|
55
|
+
4. If confidence is low or uncertainty is high, tell user what is ambiguous.
|
|
56
|
+
|
|
57
|
+
## Output Contract
|
|
58
|
+
|
|
59
|
+
- `summary`: high-level description
|
|
60
|
+
- `ocr.full_text` + `ocr.lines`: extracted text evidence
|
|
61
|
+
- `layout.regions`: structural/layout blocks with reading order
|
|
62
|
+
- `semantics`: entities, scene, intent, relations
|
|
63
|
+
- `visual`: color/style clues
|
|
64
|
+
- `uncertainty`: uncertain points
|
|
65
|
+
|
|
66
|
+
Detailed schema: `references/output-schema.md`
|
|
67
|
+
|
|
68
|
+
## Failure Handling
|
|
69
|
+
|
|
70
|
+
- If command fails due to missing auth or quota, report exact error and ask user to check backend setup (e.g., run `gemini` for Gemini CLI login).
|
|
71
|
+
- If JSON is partially malformed, keep raw text and continue with best-effort extraction.
|
|
72
|
+
|
|
73
|
+
## Implementation Note
|
|
74
|
+
|
|
75
|
+
v1 uses Gemini CLI as the default vision backend (`gemini -p` with JSON output mode). The architecture is designed to support additional vision engines (PaddleOCR, DeepSeek OCR, etc.) in future versions.
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# ModLens Output Schema (v1)
|
|
2
|
+
|
|
3
|
+
```json
|
|
4
|
+
{
|
|
5
|
+
"summary": "string",
|
|
6
|
+
"ocr": {
|
|
7
|
+
"full_text": "string",
|
|
8
|
+
"lines": [
|
|
9
|
+
{
|
|
10
|
+
"text": "string",
|
|
11
|
+
"language": "string",
|
|
12
|
+
"confidence": 0
|
|
13
|
+
}
|
|
14
|
+
]
|
|
15
|
+
},
|
|
16
|
+
"layout": {
|
|
17
|
+
"regions": [
|
|
18
|
+
{
|
|
19
|
+
"id": "string",
|
|
20
|
+
"type": "title|subtitle|paragraph|list|table|chart|form|image|icon|other",
|
|
21
|
+
"bbox": { "x": 0, "y": 0, "w": 0, "h": 0 },
|
|
22
|
+
"reading_order": 1,
|
|
23
|
+
"text": "string"
|
|
24
|
+
}
|
|
25
|
+
]
|
|
26
|
+
},
|
|
27
|
+
"semantics": {
|
|
28
|
+
"scene": "string",
|
|
29
|
+
"intent": "string",
|
|
30
|
+
"entities": [
|
|
31
|
+
{
|
|
32
|
+
"name": "string",
|
|
33
|
+
"type": "string",
|
|
34
|
+
"evidence": "string"
|
|
35
|
+
}
|
|
36
|
+
],
|
|
37
|
+
"relations": [
|
|
38
|
+
{
|
|
39
|
+
"subject": "string",
|
|
40
|
+
"predicate": "string",
|
|
41
|
+
"object": "string"
|
|
42
|
+
}
|
|
43
|
+
]
|
|
44
|
+
},
|
|
45
|
+
"visual": {
|
|
46
|
+
"dominant_colors": ["string"],
|
|
47
|
+
"style": "string",
|
|
48
|
+
"notes": ["string"]
|
|
49
|
+
},
|
|
50
|
+
"uncertainty": ["string"]
|
|
51
|
+
}
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Notes:
|
|
55
|
+
- `confidence` is 0-1 numeric estimate.
|
|
56
|
+
- `bbox` uses image-relative coordinates.
|
|
57
|
+
- If a field is unavailable, return empty string/array and explain in `uncertainty`.
|