@liustack/modlens 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Leon Liu (liustack)
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,77 @@
1
+ # ModLens
2
+
3
+ A CLI toolkit for AI agents that converts image sources (local path or remote URL) into structured text evidence, bridging the vision gap for text-only LLM workflows.
4
+
5
+ [中文说明](README.zh-CN.md)
6
+
7
+ ## Features
8
+
9
+ - Built for non-vision LLM setups (text-only models + external vision bridge)
10
+ - Supports local image paths and remote image URLs
11
+ - Pluggable vision backend — ships with Gemini CLI; more engines (PaddleOCR, DeepSeek, etc.) planned
12
+ - Outputs machine-consumable JSON (OCR + layout + semantics + visual clues)
13
+ - Designed to be called from Agent Skills (Claude Code, Codex, Cursor, etc.)
14
+
15
+ ## Install
16
+
17
+ ```bash
18
+ npm install -g @liustack/modlens
19
+ ```
20
+
21
+ The default backend requires Gemini CLI to be installed and authenticated:
22
+
23
+ ```bash
24
+ npm install -g @google/gemini-cli
25
+ gemini
26
+ ```
27
+
28
+ Or run with `npx`:
29
+
30
+ ```bash
31
+ npx @liustack/modlens [options]
32
+ ```
33
+
34
+ ## Usage
35
+
36
+ ```bash
37
+ # Print JSON result to stdout
38
+ modlens -i screenshot.png
39
+
40
+ # Save to file
41
+ modlens -i screenshot.png -o lens.json
42
+
43
+ # Specify model + extra prompt constraints
44
+ modlens -i screenshot.png -m gemini-2.5-flash --prompt "Focus on table structure"
45
+ ```
46
+
47
+ ## Options
48
+
49
+ | Flag | Description |
50
+ |------|-------------|
51
+ | `-i, --input <path>` | Input image path (required) |
52
+ | `-o, --output <path>` | Write result JSON to a file |
53
+ | `-m, --model <name>` | Vision model name (backend-specific) |
54
+ | `--prompt <text>` | Extra extraction constraints |
55
+ | `--timeout <ms>` | Timeout in milliseconds (default: `180000`) |
56
+ | `--gemini-bin <path>` | Gemini CLI binary path (default: `gemini`) |
57
+
58
+ ## Vision Backends
59
+
60
+ ModLens uses a pluggable architecture for vision recognition. The current v1 ships with **Gemini CLI** as the default backend. Future versions will support additional engines such as PaddleOCR, DeepSeek OCR, and other multimodal/vision-capable models.
61
+
62
+ ## Agent Skill
63
+
64
+ - [modlens/SKILL.md](skills/modlens/SKILL.md)
65
+
66
+ ## Notes
67
+
68
+ - `modlens` focuses on visual parsing only.
69
+ - `modsearch` and `modfetch` belong to separate projects and are intentionally out of scope.
70
+
71
+ ## Disclaimer
72
+
73
+ This project is for **personal learning and experimentation only**. It is not intended for commercial use.
74
+
75
+ ## License
76
+
77
+ MIT
@@ -0,0 +1,69 @@
1
+ # ModLens
2
+
3
+ 面向 AI Agent 的视觉外挂 CLI,用于把图片来源(本地路径或远程 URL)转成结构化文本证据,补齐纯文本 LLM 的“看图”能力。
4
+
5
+ ## 特性
6
+
7
+ - 面向“无视觉能力模型”场景(文本模型 + 外挂视觉)
8
+ - 支持本地图片路径与远程图片 URL
9
+ - 基于 Gemini CLI 非交互调用(`gemini -p`)
10
+ - 输出结构化 JSON(OCR + 布局 + 语义 + 视觉线索)
11
+ - 适合作为 Agent Skill 工具被 Claude Code / Codex 等调用
12
+
13
+ ## 安装
14
+
15
+ ```bash
16
+ npm install -g @liustack/modlens
17
+ ```
18
+
19
+ 需要先安装并认证 Gemini CLI:
20
+
21
+ ```bash
22
+ npm install -g @google/gemini-cli
23
+ gemini
24
+ ```
25
+
26
+ 或直接用 `npx`:
27
+
28
+ ```bash
29
+ npx @liustack/modlens [options]
30
+ ```
31
+
32
+ ## 用法
33
+
34
+ ```bash
35
+ # 标准输出 JSON
36
+ modlens -i screenshot.png
37
+
38
+ # 落盘到文件
39
+ modlens -i screenshot.png -o lens.json
40
+
41
+ # 指定模型和额外解析要求
42
+ modlens -i screenshot.png -m gemini-2.5-flash --prompt "重点提取表格结构"
43
+ ```
44
+
45
+ ## 参数
46
+
47
+ - `-i, --input <path>` 输入图片路径(必填)
48
+ - `-o, --output <path>` 可选输出 JSON 路径
49
+ - `-m, --model <name>` Gemini 模型名
50
+ - `--prompt <text>` 额外解析约束
51
+ - `--timeout <ms>` 超时毫秒(默认 `180000`)
52
+ - `--gemini-bin <path>` Gemini 可执行路径(默认 `gemini`)
53
+
54
+ ## Agent Skill
55
+
56
+ - [modlens/SKILL.md](skills/modlens/SKILL.md)
57
+
58
+ ## 说明
59
+
60
+ - `modlens` 只做视觉解析。
61
+ - `modsearch` / `modfetch` 是其他独立项目,不在本仓库实现。
62
+
63
+ ## 免责声明
64
+
65
+ 本项目仅供**个人学习与实验**使用,请勿用于商业用途。
66
+
67
+ ## License
68
+
69
+ MIT
package/dist/main.js ADDED
@@ -0,0 +1,275 @@
1
+ #!/usr/bin/env node
2
+ import { Command } from "commander";
3
+ import * as fs from "fs";
4
+ import * as path from "path";
5
+ import { spawn } from "child_process";
6
+ function buildVisionPrompt(imageAbsolutePath, extraPrompt) {
7
+ const basePrompt = `Analyze this image: ${imageAbsolutePath}
8
+
9
+ You are an image parsing engine.
10
+ Analyze the provided image.
11
+
12
+ Goal: Convert all image information into structured results consumable by a text-only LLM.
13
+
14
+ Strict requirements:
15
+ 1. Output JSON only. No Markdown, no explanatory text.
16
+ 2. Cover all visible text, semantics, structure, layout, and visual clues as thoroughly as possible.
17
+ 3. If any information is uncertain, note it in the uncertainty field.
18
+
19
+ Output JSON structure:
20
+ {
21
+ "summary": "",
22
+ "ocr": {
23
+ "full_text": "",
24
+ "lines": [
25
+ {
26
+ "text": "",
27
+ "language": "",
28
+ "confidence": 0
29
+ }
30
+ ]
31
+ },
32
+ "layout": {
33
+ "regions": [
34
+ {
35
+ "id": "",
36
+ "type": "title|subtitle|paragraph|list|table|chart|form|image|icon|other",
37
+ "bbox": { "x": 0, "y": 0, "w": 0, "h": 0 },
38
+ "reading_order": 1,
39
+ "text": ""
40
+ }
41
+ ]
42
+ },
43
+ "semantics": {
44
+ "scene": "",
45
+ "intent": "",
46
+ "entities": [
47
+ {
48
+ "name": "",
49
+ "type": "",
50
+ "evidence": ""
51
+ }
52
+ ],
53
+ "relations": [
54
+ {
55
+ "subject": "",
56
+ "predicate": "",
57
+ "object": ""
58
+ }
59
+ ]
60
+ },
61
+ "visual": {
62
+ "dominant_colors": [""],
63
+ "style": "",
64
+ "notes": [""]
65
+ },
66
+ "uncertainty": [""]
67
+ }`;
68
+ if (!extraPrompt || !extraPrompt.trim()) {
69
+ return basePrompt;
70
+ }
71
+ return `${basePrompt}
72
+
73
+ Additional requirements:
74
+ ${extraPrompt.trim()}`;
75
+ }
76
+ const DEFAULT_TIMEOUT_MS = 18e4;
77
+ function buildGeminiInvocation(options) {
78
+ const isRemote = isRemoteSource(options.imagePath);
79
+ const imageSource = isRemote ? options.imagePath.trim() : path.resolve(options.imagePath);
80
+ const imageDir = path.dirname(imageSource);
81
+ const prompt = buildVisionPrompt(imageSource, options.extraPrompt);
82
+ const args = ["-p", prompt, "--output-format", "json"];
83
+ if (options.model) {
84
+ args.push("-m", options.model);
85
+ }
86
+ return {
87
+ command: options.geminiBin || "gemini",
88
+ args,
89
+ cwd: options.workspaceDir || (isRemote ? process.cwd() : imageDir)
90
+ };
91
+ }
92
+ function parseGeminiCliJsonOutput(stdout) {
93
+ let parsed;
94
+ try {
95
+ parsed = JSON.parse(stdout.trim());
96
+ } catch (error) {
97
+ throw new Error(`Failed to parse Gemini JSON output: ${error.message}`);
98
+ }
99
+ if (!parsed || typeof parsed !== "object" || typeof parsed.response !== "string") {
100
+ throw new Error("Gemini JSON output is missing a string `response` field.");
101
+ }
102
+ return parsed;
103
+ }
104
+ function extractStructuredResponse(text) {
105
+ const rawText = text.trim();
106
+ const direct = tryParseJson(rawText);
107
+ if (direct !== null) {
108
+ return {
109
+ structured: direct,
110
+ rawText
111
+ };
112
+ }
113
+ const fencedMatch = /```(?:json)?\s*([\s\S]*?)```/i.exec(rawText);
114
+ if (fencedMatch) {
115
+ const parsedFenced = tryParseJson(fencedMatch[1].trim());
116
+ if (parsedFenced !== null) {
117
+ return {
118
+ structured: parsedFenced,
119
+ rawText
120
+ };
121
+ }
122
+ }
123
+ const firstBrace = rawText.indexOf("{");
124
+ const lastBrace = rawText.lastIndexOf("}");
125
+ if (firstBrace >= 0 && lastBrace > firstBrace) {
126
+ const possibleJson = rawText.slice(firstBrace, lastBrace + 1);
127
+ const parsedObject = tryParseJson(possibleJson);
128
+ if (parsedObject !== null) {
129
+ return {
130
+ structured: parsedObject,
131
+ rawText
132
+ };
133
+ }
134
+ }
135
+ return {
136
+ structured: null,
137
+ rawText
138
+ };
139
+ }
140
+ async function analyzeImage(options) {
141
+ const resolvedInput = resolveInput(options.input);
142
+ if (resolvedInput.kind === "local") {
143
+ validateInputFile(resolvedInput.source);
144
+ }
145
+ const invocation = buildGeminiInvocation({
146
+ imagePath: resolvedInput.source,
147
+ model: options.model,
148
+ geminiBin: options.geminiBin,
149
+ workspaceDir: options.workspaceDir,
150
+ extraPrompt: options.prompt
151
+ });
152
+ const commandResult = await runCommand(invocation, options.timeoutMs ?? DEFAULT_TIMEOUT_MS);
153
+ const geminiOutput = parseGeminiCliJsonOutput(commandResult.stdout);
154
+ const extracted = extractStructuredResponse(geminiOutput.response);
155
+ return {
156
+ image: resolvedInput.source,
157
+ structured: extracted.structured,
158
+ rawText: extracted.rawText,
159
+ meta: {
160
+ generatedAt: (/* @__PURE__ */ new Date()).toISOString(),
161
+ model: options.model ?? null,
162
+ geminiSessionId: geminiOutput.session_id ?? null,
163
+ geminiStats: geminiOutput.stats ?? null
164
+ }
165
+ };
166
+ }
167
+ function resolveInput(input) {
168
+ const trimmed = input.trim();
169
+ if (!trimmed) {
170
+ throw new Error("Input path is required.");
171
+ }
172
+ if (isRemoteSource(trimmed)) {
173
+ return { source: trimmed, kind: "remote" };
174
+ }
175
+ if (/^file:\/\//i.test(trimmed)) {
176
+ const localPath = decodeURI(trimmed.replace(/^file:\/\//i, ""));
177
+ return { source: path.resolve(localPath), kind: "local" };
178
+ }
179
+ return { source: path.resolve(trimmed), kind: "local" };
180
+ }
181
+ function isRemoteSource(value) {
182
+ return /^https?:\/\//i.test(value.trim());
183
+ }
184
+ function validateInputFile(filePath) {
185
+ if (!fs.existsSync(filePath)) {
186
+ throw new Error(`Input image not found: ${filePath}`);
187
+ }
188
+ const stat = fs.statSync(filePath);
189
+ if (!stat.isFile()) {
190
+ throw new Error(`Input is not a file: ${filePath}`);
191
+ }
192
+ }
193
+ function tryParseJson(text) {
194
+ try {
195
+ return JSON.parse(text);
196
+ } catch {
197
+ return null;
198
+ }
199
+ }
200
+ function runCommand(invocation, timeoutMs) {
201
+ return new Promise((resolve, reject) => {
202
+ const child = spawn(invocation.command, invocation.args, {
203
+ cwd: invocation.cwd,
204
+ stdio: ["ignore", "pipe", "pipe"]
205
+ });
206
+ let stdout = "";
207
+ let stderr = "";
208
+ let timedOut = false;
209
+ const timer = setTimeout(() => {
210
+ timedOut = true;
211
+ child.kill("SIGTERM");
212
+ }, timeoutMs);
213
+ child.stdout.on("data", (chunk) => {
214
+ stdout += chunk.toString();
215
+ });
216
+ child.stderr.on("data", (chunk) => {
217
+ stderr += chunk.toString();
218
+ });
219
+ child.on("error", (error) => {
220
+ clearTimeout(timer);
221
+ if (error.code === "ENOENT") {
222
+ reject(new Error(`Gemini CLI not found: ${invocation.command}`));
223
+ return;
224
+ }
225
+ reject(error);
226
+ });
227
+ child.on("close", (code) => {
228
+ clearTimeout(timer);
229
+ if (timedOut) {
230
+ reject(new Error(`Gemini CLI timed out after ${timeoutMs} ms.`));
231
+ return;
232
+ }
233
+ if (code !== 0) {
234
+ reject(
235
+ new Error(
236
+ `Gemini CLI failed with code ${code}.${stderr ? ` stderr: ${stderr.trim()}` : ""}`
237
+ )
238
+ );
239
+ return;
240
+ }
241
+ resolve({ stdout, stderr });
242
+ });
243
+ });
244
+ }
245
+ const program = new Command();
246
+ program.name("modlens").description("Image-to-text visual bridge for non-vision LLM workflows").version("0.1.0").requiredOption("-i, --input <path>", "Input image path").option("-o, --output <path>", "Write result JSON to a file").option("-m, --model <name>", "Gemini model name").option("--prompt <text>", "Extra prompt constraints for this image").option("--timeout <ms>", "Command timeout in milliseconds", "180000").option("--gemini-bin <path>", "Gemini CLI binary path", "gemini").action(async (options) => {
247
+ try {
248
+ const timeoutMs = Number.parseInt(options.timeout, 10);
249
+ if (!Number.isFinite(timeoutMs) || timeoutMs <= 0) {
250
+ throw new Error("Invalid --timeout. Use a positive integer in milliseconds.");
251
+ }
252
+ const result = await analyzeImage({
253
+ input: options.input,
254
+ model: options.model,
255
+ prompt: options.prompt,
256
+ timeoutMs,
257
+ geminiBin: options.geminiBin
258
+ });
259
+ const output = JSON.stringify(result, null, 2);
260
+ if (options.output) {
261
+ const outputPath = path.resolve(options.output);
262
+ fs.mkdirSync(path.dirname(outputPath), { recursive: true });
263
+ fs.writeFileSync(outputPath, output, "utf-8");
264
+ }
265
+ process.stdout.write(`${output}
266
+ `);
267
+ } catch (error) {
268
+ process.stderr.write(
269
+ `Error: ${error instanceof Error ? error.message : String(error)}
270
+ `
271
+ );
272
+ process.exit(1);
273
+ }
274
+ });
275
+ program.parse();
package/package.json ADDED
@@ -0,0 +1,54 @@
1
+ {
2
+ "name": "@liustack/modlens",
3
+ "version": "0.1.0",
4
+ "description": "CLI tool to provide visual understanding for non-vision LLMs",
5
+ "type": "module",
6
+ "bin": {
7
+ "modlens": "./dist/main.js"
8
+ },
9
+ "files": [
10
+ "dist",
11
+ "skills/modlens/SKILL.md",
12
+ "skills/modlens/references"
13
+ ],
14
+ "keywords": [
15
+ "cli",
16
+ "vision",
17
+ "ocr",
18
+ "gemini",
19
+ "claude-code",
20
+ "skill",
21
+ "image-to-text",
22
+ "multimodal",
23
+ "modlens"
24
+ ],
25
+ "author": "Leon Liu",
26
+ "license": "MIT",
27
+ "repository": {
28
+ "type": "git",
29
+ "url": "git+https://github.com/liustack/modlens.git"
30
+ },
31
+ "bugs": {
32
+ "url": "https://github.com/liustack/modlens/issues"
33
+ },
34
+ "homepage": "https://github.com/liustack/modlens#readme",
35
+ "engines": {
36
+ "node": ">=18"
37
+ },
38
+ "dependencies": {
39
+ "commander": "^13.1.0"
40
+ },
41
+ "devDependencies": {
42
+ "@types/node": "^22.19.7",
43
+ "typescript": "^5.9.3",
44
+ "vite": "^6.4.1",
45
+ "vitest": "^3.2.4"
46
+ },
47
+ "scripts": {
48
+ "dev": "vite build --watch",
49
+ "build": "vite build",
50
+ "typecheck": "tsc --noEmit",
51
+ "test": "vitest run",
52
+ "docs:list": "node scripts/docs-list.js"
53
+ }
54
+ }
@@ -0,0 +1,75 @@
1
+ ---
2
+ name: modlens
3
+ description: "Bridge image understanding for non-vision LLM workflows. Use when user provides screenshots/photos/charts/doc images and the active model lacks multimodal vision. Call modlens to extract OCR text, semantics, structure, and layout as JSON evidence."
4
+ allowed-tools:
5
+ - Bash
6
+ ---
7
+
8
+ # ModLens — Vision Bridge Skill
9
+
10
+ Use this skill when:
11
+ - User asks to analyze an image/screenshot/chart/document photo
12
+ - Current model cannot directly read images
13
+ - You need structured visual evidence before downstream reasoning
14
+
15
+ Do not use this skill for:
16
+ - Web search (`modsearch`)
17
+ - Web fetch (`modfetch`)
18
+
19
+ ## Prerequisites
20
+
21
+ ```bash
22
+ modlens --version
23
+ ```
24
+
25
+ The default vision backend (Gemini CLI) must also be installed and authenticated:
26
+
27
+ ```bash
28
+ gemini --version
29
+ ```
30
+
31
+ If `gemini` is missing:
32
+
33
+ ```bash
34
+ npm install -g @google/gemini-cli
35
+ gemini
36
+ ```
37
+
38
+ ## Command
39
+
40
+ ```bash
41
+ modlens -i <image-path>
42
+ ```
43
+
44
+ Optional:
45
+
46
+ ```bash
47
+ modlens -i <image-path> -o <output-json-path> -m <model-name> --prompt "<extra constraints>"
48
+ ```
49
+
50
+ ## Workflow
51
+
52
+ 1. If user includes one or more images, run `modlens` for each image.
53
+ 2. Parse returned JSON.
54
+ 3. Feed `summary`, `ocr`, `layout`, and `semantics` back into your reasoning context.
55
+ 4. If confidence is low or uncertainty is high, tell user what is ambiguous.
56
+
57
+ ## Output Contract
58
+
59
+ - `summary`: high-level description
60
+ - `ocr.full_text` + `ocr.lines`: extracted text evidence
61
+ - `layout.regions`: structural/layout blocks with reading order
62
+ - `semantics`: entities, scene, intent, relations
63
+ - `visual`: color/style clues
64
+ - `uncertainty`: uncertain points
65
+
66
+ Detailed schema: `references/output-schema.md`
67
+
68
+ ## Failure Handling
69
+
70
+ - If command fails due to missing auth or quota, report exact error and ask user to check backend setup (e.g., run `gemini` for Gemini CLI login).
71
+ - If JSON is partially malformed, keep raw text and continue with best-effort extraction.
72
+
73
+ ## Implementation Note
74
+
75
+ v1 uses Gemini CLI as the default vision backend (`gemini -p` with JSON output mode). The architecture is designed to support additional vision engines (PaddleOCR, DeepSeek OCR, etc.) in future versions.
@@ -0,0 +1,57 @@
1
+ # ModLens Output Schema (v1)
2
+
3
+ ```json
4
+ {
5
+ "summary": "string",
6
+ "ocr": {
7
+ "full_text": "string",
8
+ "lines": [
9
+ {
10
+ "text": "string",
11
+ "language": "string",
12
+ "confidence": 0
13
+ }
14
+ ]
15
+ },
16
+ "layout": {
17
+ "regions": [
18
+ {
19
+ "id": "string",
20
+ "type": "title|subtitle|paragraph|list|table|chart|form|image|icon|other",
21
+ "bbox": { "x": 0, "y": 0, "w": 0, "h": 0 },
22
+ "reading_order": 1,
23
+ "text": "string"
24
+ }
25
+ ]
26
+ },
27
+ "semantics": {
28
+ "scene": "string",
29
+ "intent": "string",
30
+ "entities": [
31
+ {
32
+ "name": "string",
33
+ "type": "string",
34
+ "evidence": "string"
35
+ }
36
+ ],
37
+ "relations": [
38
+ {
39
+ "subject": "string",
40
+ "predicate": "string",
41
+ "object": "string"
42
+ }
43
+ ]
44
+ },
45
+ "visual": {
46
+ "dominant_colors": ["string"],
47
+ "style": "string",
48
+ "notes": ["string"]
49
+ },
50
+ "uncertainty": ["string"]
51
+ }
52
+ ```
53
+
54
+ Notes:
55
+ - `confidence` is 0-1 numeric estimate.
56
+ - `bbox` uses image-relative coordinates.
57
+ - If a field is unavailable, return empty string/array and explain in `uncertainty`.