@liustack/modlens 2.2.0 → 2.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +10 -0
- package/README.zh-CN.md +10 -0
- package/dist/main.js +147 -1
- package/package.json +1 -1
- package/skills/modlens/SKILL.md +1 -0
package/README.md
CHANGED
|
@@ -90,6 +90,10 @@ Batch mode works too: drop three illustrations at once, and the model announces
|
|
|
90
90
|
|
|
91
91
|

|
|
92
92
|
|
|
93
|
+
Stress test: a scatter plot of 128 models. ModLens pulls out the axes, the log scale, and the highlighted DeepSeek V4 Flash point at $0.028 and score 50, then walks through the cost-performance cutoff line. Dense charts are where vision models usually fold; this one holds.
|
|
94
|
+
|
|
95
|
+

|
|
96
|
+
|
|
93
97
|
## CLI reference
|
|
94
98
|
|
|
95
99
|
```bash
|
|
@@ -141,6 +145,12 @@ One catch: once text-only is declared, the Codex TUI **blocks Ctrl+V image paste
|
|
|
141
145
|
- **Drag the image file into the terminal**, or type its path. The path lands as plain text, and the modlens skill picks it up from there.
|
|
142
146
|
- Attach it with `codex exec -i image.png "..."`. The skill reads the path out of the message tag.
|
|
143
147
|
|
|
148
|
+
## Using it in Claude Code (gateway models)
|
|
149
|
+
|
|
150
|
+
No setup needed: drag the image file into the terminal, or type its path, and the skill takes over.
|
|
151
|
+
|
|
152
|
+
Paste is trickier. If you run a text-only model behind `ANTHROPIC_BASE_URL`, Claude Code never writes pasted images to a regular temp file and has no modality switch, so a pasted image reaches the model as a pathless `[Unsupported Image]` placeholder (lenient gateways like DeepSeek's Anthropic endpoint) or breaks the request outright ([#62009](https://github.com/anthropics/claude-code/issues/62009)). But the bytes are not gone: Claude Code appends every user message, images included, to the local session transcript before the gateway ever sees it. That is what `modlens recover-paste` exploits: it pulls the most recent pasted images back out of the transcript and prints real file paths, ready for `modlens -i`. The skill runs this automatically when it spots the placeholder. Transcripts are per-session files; recovery picks the one holding the newest pasted image by message timestamp, so concurrent sessions in the same project do not confuse it. One honest caveat: the transcript layout is Claude Code internals with no compatibility promise; if recovery ever breaks, dragging the file still works everywhere.
|
|
153
|
+
|
|
144
154
|
## Why a bridge instead of a multimodal model?
|
|
145
155
|
|
|
146
156
|
- **Keep your model.** You picked DeepSeek-V4-Flash (or gpt-oss, or whatever else) for its price and its reasoning, not its eyesight. ModLens adds sight without touching that choice.
|
package/README.zh-CN.md
CHANGED
|
@@ -90,6 +90,10 @@ npx @liustack/modlens -i workflow.jpg
|
|
|
90
90
|
|
|
91
91
|

|
|
92
92
|
|
|
93
|
+
压力测试:一张 128 个模型的智能对成本散点图。ModLens 读出双轴、对数刻度,把高亮的 DeepSeek V4 Flash 精准拎出来(成本约 $0.028、智能指数 50),还讲明白了性价比斩杀线。密集图表是识图模型最容易露怯的地方,这一关它扛住了。
|
|
94
|
+
|
|
95
|
+

|
|
96
|
+
|
|
93
97
|
## CLI 参数
|
|
94
98
|
|
|
95
99
|
```bash
|
|
@@ -141,6 +145,12 @@ Codex 只认 Responses API,DeepSeek 官方端点原生支持。先照着[官
|
|
|
141
145
|
- **把图片文件拖进终端**,或者手打路径。路径以纯文本形式落进消息,modlens skill 接着从这里接手。
|
|
142
146
|
- 用 `codex exec -i 图片.png "..."` skill 从这里把路径抠出来。
|
|
143
147
|
|
|
148
|
+
## 在 Claude Code 里用(网关接第三方模型)
|
|
149
|
+
|
|
150
|
+
不用任何配置:把图片文件拖进终端,或手打路径,skill 直接接手。
|
|
151
|
+
|
|
152
|
+
粘贴要多说两句。走 `ANTHROPIC_BASE_URL` 网关跑纯文本模型时,Claude Code 粘贴的图片从不写普通临时文件,也没有声明模型无视觉的开关,粘贴的图要么变成一个不带路径的 `[Unsupported Image]` 占位符到达模型(DeepSeek 的 Anthropic 兼容端点这类宽容网关),要么直接把请求搞挂([#62009](https://github.com/anthropics/claude-code/issues/62009))。但图片字节没有蒸发:Claude Code 在网关看到消息之前,就把每条用户消息(含图片)原样写进了本地会话记录。`modlens recover-paste` 干的就是这件事:从会话记录里把最近粘贴的图捞回来,落成真实文件路径,直接喂给 `modlens -i`。skill 看到占位符会自动跑这一步。会话记录本来就是一个会话一个文件,recover-paste 按消息时间戳挑「持有最新粘贴图」的那份,同项目并发开多个会话也不会拿错。一句老实话:会话记录格式是 Claude Code 的内部实现,没有兼容承诺,哪天捞不动了,拖文件永远是保底。
|
|
153
|
+
|
|
144
154
|
## 为什么外挂,而不是换多模态模型?
|
|
145
155
|
|
|
146
156
|
- **模型不用换。** 你选 DeepSeek-V4-Flash(或 gpt-oss,或别的什么)图的是价格和推理能力,不是视力。ModLens 只加视力,不碰这个选择。
|
package/dist/main.js
CHANGED
|
@@ -4,6 +4,7 @@ import * as fs from "fs";
|
|
|
4
4
|
import * as path from "path";
|
|
5
5
|
import { spawn } from "child_process";
|
|
6
6
|
import * as os from "os";
|
|
7
|
+
import * as crypto from "crypto";
|
|
7
8
|
const CONFIG_DIR = path.join(os.homedir(), ".modlens");
|
|
8
9
|
const CONFIG_PATH = path.join(CONFIG_DIR, "config.json");
|
|
9
10
|
const ENV_BINDINGS = {
|
|
@@ -843,8 +844,129 @@ function runCommand(providerName, invocation, timeoutMs) {
|
|
|
843
844
|
});
|
|
844
845
|
});
|
|
845
846
|
}
|
|
847
|
+
const EXT_BY_MIME = {
|
|
848
|
+
"image/png": "png",
|
|
849
|
+
"image/jpeg": "jpg",
|
|
850
|
+
"image/webp": "webp",
|
|
851
|
+
"image/gif": "gif"
|
|
852
|
+
};
|
|
853
|
+
function projectSlug(cwd) {
|
|
854
|
+
return path.resolve(cwd).replace(/[/.]/g, "-");
|
|
855
|
+
}
|
|
856
|
+
function locateTranscript(cwd) {
|
|
857
|
+
const dir = path.join(os.homedir(), ".claude", "projects", projectSlug(cwd));
|
|
858
|
+
let entries;
|
|
859
|
+
try {
|
|
860
|
+
entries = fs.readdirSync(dir).filter((name) => name.endsWith(".jsonl"));
|
|
861
|
+
} catch {
|
|
862
|
+
throw new Error(
|
|
863
|
+
`No Claude Code transcripts found for this directory (${dir}). Run from the project the image was pasted in, or pass --transcript <path>.`
|
|
864
|
+
);
|
|
865
|
+
}
|
|
866
|
+
if (entries.length === 0) {
|
|
867
|
+
throw new Error(`No transcripts in ${dir}. Pass --transcript <path> to pick one manually.`);
|
|
868
|
+
}
|
|
869
|
+
let best = null;
|
|
870
|
+
for (const name of entries) {
|
|
871
|
+
const full = path.join(dir, name);
|
|
872
|
+
const timestamp = lastImageTimestamp(full);
|
|
873
|
+
if (timestamp && (!best || timestamp > best.timestamp)) {
|
|
874
|
+
best = { full, timestamp };
|
|
875
|
+
}
|
|
876
|
+
}
|
|
877
|
+
if (!best) {
|
|
878
|
+
throw new Error(
|
|
879
|
+
`No pasted images found in any transcript under ${dir}. The user may not have pasted any, or the transcript format changed; ask for a file path instead.`
|
|
880
|
+
);
|
|
881
|
+
}
|
|
882
|
+
return best.full;
|
|
883
|
+
}
|
|
884
|
+
function lastImageTimestamp(transcriptPath) {
|
|
885
|
+
let raw;
|
|
886
|
+
try {
|
|
887
|
+
raw = fs.readFileSync(transcriptPath, "utf-8");
|
|
888
|
+
} catch {
|
|
889
|
+
return null;
|
|
890
|
+
}
|
|
891
|
+
let latest = null;
|
|
892
|
+
for (const line of raw.split("\n")) {
|
|
893
|
+
if (!line.includes('"image"')) {
|
|
894
|
+
continue;
|
|
895
|
+
}
|
|
896
|
+
let parsed;
|
|
897
|
+
try {
|
|
898
|
+
parsed = JSON.parse(line);
|
|
899
|
+
} catch {
|
|
900
|
+
continue;
|
|
901
|
+
}
|
|
902
|
+
const entry = parsed;
|
|
903
|
+
if (entry.message?.role !== "user" || !Array.isArray(entry.message.content)) {
|
|
904
|
+
continue;
|
|
905
|
+
}
|
|
906
|
+
const hasImage = entry.message.content.some(
|
|
907
|
+
(block) => block?.type === "image" && block.source?.type === "base64"
|
|
908
|
+
);
|
|
909
|
+
if (hasImage && entry.timestamp && (!latest || entry.timestamp > latest)) {
|
|
910
|
+
latest = entry.timestamp;
|
|
911
|
+
}
|
|
912
|
+
}
|
|
913
|
+
return latest;
|
|
914
|
+
}
|
|
915
|
+
function extractUserImages(transcriptPath) {
|
|
916
|
+
let raw;
|
|
917
|
+
try {
|
|
918
|
+
raw = fs.readFileSync(transcriptPath, "utf-8");
|
|
919
|
+
} catch (error) {
|
|
920
|
+
throw new Error(`Cannot read transcript ${transcriptPath}: ${error.message}`);
|
|
921
|
+
}
|
|
922
|
+
const images = [];
|
|
923
|
+
for (const line of raw.split("\n")) {
|
|
924
|
+
if (!line.includes('"image"')) {
|
|
925
|
+
continue;
|
|
926
|
+
}
|
|
927
|
+
let parsed;
|
|
928
|
+
try {
|
|
929
|
+
parsed = JSON.parse(line);
|
|
930
|
+
} catch {
|
|
931
|
+
continue;
|
|
932
|
+
}
|
|
933
|
+
const message = parsed.message;
|
|
934
|
+
if (message?.role !== "user" || !Array.isArray(message.content)) {
|
|
935
|
+
continue;
|
|
936
|
+
}
|
|
937
|
+
for (const block of message.content) {
|
|
938
|
+
const source = block?.source;
|
|
939
|
+
if (block?.type === "image" && source?.type === "base64" && source.data) {
|
|
940
|
+
images.push({ mediaType: source.media_type ?? "image/png", data: source.data });
|
|
941
|
+
}
|
|
942
|
+
}
|
|
943
|
+
}
|
|
944
|
+
return images;
|
|
945
|
+
}
|
|
946
|
+
function recoverPastedImages(options = {}) {
|
|
947
|
+
const transcript = options.transcript ?? locateTranscript(options.cwd ?? process.cwd());
|
|
948
|
+
const count = Math.max(1, options.count ?? 1);
|
|
949
|
+
const outDir = options.outDir ?? path.join(os.tmpdir(), "modlens-paste");
|
|
950
|
+
const all = extractUserImages(transcript);
|
|
951
|
+
if (all.length === 0) {
|
|
952
|
+
throw new Error(
|
|
953
|
+
`No pasted images found in ${transcript}. The user may not have pasted any, or the transcript format changed; ask for a file path instead.`
|
|
954
|
+
);
|
|
955
|
+
}
|
|
956
|
+
fs.mkdirSync(outDir, { recursive: true });
|
|
957
|
+
const picked = all.slice(-count);
|
|
958
|
+
const images = picked.map((image) => {
|
|
959
|
+
const buffer = Buffer.from(image.data, "base64");
|
|
960
|
+
const hash = crypto.createHash("sha256").update(buffer).digest("hex").slice(0, 8);
|
|
961
|
+
const ext = EXT_BY_MIME[image.mediaType] ?? "png";
|
|
962
|
+
const filePath = path.join(outDir, `paste-${hash}.${ext}`);
|
|
963
|
+
fs.writeFileSync(filePath, buffer);
|
|
964
|
+
return { path: filePath, mediaType: image.mediaType, bytes: buffer.length };
|
|
965
|
+
});
|
|
966
|
+
return { transcript, images };
|
|
967
|
+
}
|
|
846
968
|
const program = new Command();
|
|
847
|
-
program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("2.
|
|
969
|
+
program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("2.3.1");
|
|
848
970
|
program.command("analyze", { isDefault: true }).description("Analyze an image into structured JSON evidence (default command)").requiredOption("-i, --input <path|url>", "Input image path or https URL").option("-o, --output <path>", "Write result JSON to a file").option("-m, --model <name>", "Provider model name").option("-p, --provider <name>", `Vision provider (${listProviders().join(", ")})`).option("--prompt <text>", "Extra focus for this image").option("--timeout <ms>", "Provider timeout in milliseconds", "180000").option("--provider-bin <path>", "Provider binary path (default: agy)").option("--workdir <path>", "Working directory for the provider").action(async (options) => {
|
|
849
971
|
try {
|
|
850
972
|
const timeoutMs = Number.parseInt(options.timeout, 10);
|
|
@@ -876,6 +998,30 @@ program.command("analyze", { isDefault: true }).description("Analyze an image in
|
|
|
876
998
|
process.exit(1);
|
|
877
999
|
}
|
|
878
1000
|
});
|
|
1001
|
+
program.command("recover-paste").description(
|
|
1002
|
+
"Recover images pasted into Claude Code from the session transcript (they never hit disk otherwise)"
|
|
1003
|
+
).option("--count <n>", "How many recent pasted images to recover", "1").option("--out-dir <path>", "Directory to write recovered images to").option("--transcript <path>", "Explicit transcript .jsonl (default: newest for cwd)").option("--cwd <path>", "Project directory the image was pasted in", process.cwd()).action(async (options) => {
|
|
1004
|
+
try {
|
|
1005
|
+
const count = Number.parseInt(options.count, 10);
|
|
1006
|
+
if (!Number.isFinite(count) || count <= 0) {
|
|
1007
|
+
throw new Error("Invalid --count. Use a positive integer.");
|
|
1008
|
+
}
|
|
1009
|
+
const result = recoverPastedImages({
|
|
1010
|
+
count,
|
|
1011
|
+
outDir: options.outDir,
|
|
1012
|
+
transcript: options.transcript,
|
|
1013
|
+
cwd: options.cwd
|
|
1014
|
+
});
|
|
1015
|
+
process.stdout.write(`${JSON.stringify(result, null, 2)}
|
|
1016
|
+
`);
|
|
1017
|
+
} catch (error) {
|
|
1018
|
+
process.stderr.write(
|
|
1019
|
+
`Error: ${error instanceof Error ? error.message : String(error)}
|
|
1020
|
+
`
|
|
1021
|
+
);
|
|
1022
|
+
process.exit(1);
|
|
1023
|
+
}
|
|
1024
|
+
});
|
|
879
1025
|
const config = program.command("config").description(`Manage ${CONFIG_PATH} (providers, keys, models)`);
|
|
880
1026
|
config.command("init").description(`Create a starter config at ${CONFIG_PATH}`).option("--force", "Overwrite an existing config file").action((options) => {
|
|
881
1027
|
try {
|
package/package.json
CHANGED
package/skills/modlens/SKILL.md
CHANGED
|
@@ -66,6 +66,7 @@ Harnesses rarely hand you a clean path. Look for these signals:
|
|
|
66
66
|
- Codex wraps every pasted or attached image in a text tag like
|
|
67
67
|
`<image name=[Image #1] path="/tmp/xxxx.png">`. Extract the `path` value and run modlens on it. Pasted images live in a temp file the harness already created.
|
|
68
68
|
- A placeholder like `image content omitted because you do not support image input` means the harness stripped an image for you. The path tag next to it still holds the real file. Use it.
|
|
69
|
+
- Claude Code never writes pasted images to a regular temp file. Behind a text-only gateway you will only see a placeholder like `[Unsupported Image]` or `[Image #1]`, with no path. When that happens, run `modlens recover-paste` (add `--count <n>` for several images): it pulls the pasted image bytes out of the local session transcript and prints the recovered file paths as JSON. Feed that path to `modlens -i`. If recovery fails (transcript format is Claude Code internals and may change), fall back to asking the user to drag the image file into the terminal or type its path.
|
|
69
70
|
- If the user mentions an image but no tag or path appears anywhere in the message, ask for the file path instead of guessing.
|
|
70
71
|
|
|
71
72
|
## Workflow
|