@liustack/modlens 2.3.0 → 2.3.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/README.zh-CN.md +1 -1
- package/dist/main.js +45 -5
- package/package.json +1 -1
- package/skills/modlens/SKILL.md +2 -1
package/README.md
CHANGED
|
@@ -149,7 +149,7 @@ One catch: once text-only is declared, the Codex TUI **blocks Ctrl+V image paste
|
|
|
149
149
|
|
|
150
150
|
No setup needed: drag the image file into the terminal, or type its path, and the skill takes over.
|
|
151
151
|
|
|
152
|
-
Paste is trickier. If you run a text-only model behind `ANTHROPIC_BASE_URL`, Claude Code never writes pasted images to a regular temp file and has no modality switch, so a pasted image reaches the model as a pathless `[Unsupported Image]` placeholder (lenient gateways like DeepSeek's Anthropic endpoint) or breaks the request outright ([#62009](https://github.com/anthropics/claude-code/issues/62009)). But the bytes are not gone: Claude Code appends every user message, images included, to the local session transcript before the gateway ever sees it. That is what `modlens recover-paste` exploits: it pulls the most recent pasted images back out of the transcript and prints real file paths, ready for `modlens -i`. The skill runs this automatically when it spots the placeholder. One honest caveat: the transcript layout is Claude Code internals with no compatibility promise; if recovery ever breaks, dragging the file still works everywhere.
|
|
152
|
+
Paste is trickier. If you run a text-only model behind `ANTHROPIC_BASE_URL`, Claude Code never writes pasted images to a regular temp file and has no modality switch, so a pasted image reaches the model as a pathless `[Unsupported Image]` placeholder (lenient gateways like DeepSeek's Anthropic endpoint) or breaks the request outright ([#62009](https://github.com/anthropics/claude-code/issues/62009)). But the bytes are not gone: Claude Code appends every user message, images included, to the local session transcript before the gateway ever sees it. That is what `modlens recover-paste` exploits: it pulls the most recent pasted images back out of the transcript and prints real file paths, ready for `modlens -i`. The skill runs this automatically when it spots the placeholder. Transcripts are per-session files; recovery picks the one holding the newest pasted image by message timestamp, so concurrent sessions in the same project do not confuse it. One honest caveat: the transcript layout is Claude Code internals with no compatibility promise; if recovery ever breaks, dragging the file still works everywhere.
|
|
153
153
|
|
|
154
154
|
## Why a bridge instead of a multimodal model?
|
|
155
155
|
|
package/README.zh-CN.md
CHANGED
|
@@ -149,7 +149,7 @@ Codex 只认 Responses API,DeepSeek 官方端点原生支持。先照着[官
|
|
|
149
149
|
|
|
150
150
|
不用任何配置:把图片文件拖进终端,或手打路径,skill 直接接手。
|
|
151
151
|
|
|
152
|
-
粘贴要多说两句。走 `ANTHROPIC_BASE_URL` 网关跑纯文本模型时,Claude Code 粘贴的图片从不写普通临时文件,也没有声明模型无视觉的开关,粘贴的图要么变成一个不带路径的 `[Unsupported Image]` 占位符到达模型(DeepSeek 的 Anthropic 兼容端点这类宽容网关),要么直接把请求搞挂([#62009](https://github.com/anthropics/claude-code/issues/62009))。但图片字节没有蒸发:Claude Code 在网关看到消息之前,就把每条用户消息(含图片)原样写进了本地会话记录。`modlens recover-paste` 干的就是这件事:从会话记录里把最近粘贴的图捞回来,落成真实文件路径,直接喂给 `modlens -i`。skill
|
|
152
|
+
粘贴要多说两句。走 `ANTHROPIC_BASE_URL` 网关跑纯文本模型时,Claude Code 粘贴的图片从不写普通临时文件,也没有声明模型无视觉的开关,粘贴的图要么变成一个不带路径的 `[Unsupported Image]` 占位符到达模型(DeepSeek 的 Anthropic 兼容端点这类宽容网关),要么直接把请求搞挂([#62009](https://github.com/anthropics/claude-code/issues/62009))。但图片字节没有蒸发:Claude Code 在网关看到消息之前,就把每条用户消息(含图片)原样写进了本地会话记录。`modlens recover-paste` 干的就是这件事:从会话记录里把最近粘贴的图捞回来,落成真实文件路径,直接喂给 `modlens -i`。skill 看到占位符会自动跑这一步。会话记录本来就是一个会话一个文件,recover-paste 按消息时间戳挑「持有最新粘贴图」的那份,同项目并发开多个会话也不会拿错。一句老实话:会话记录格式是 Claude Code 的内部实现,没有兼容承诺,哪天捞不动了,拖文件永远是保底。
|
|
153
153
|
|
|
154
154
|
## 为什么外挂,而不是换多模态模型?
|
|
155
155
|
|
package/dist/main.js
CHANGED
|
@@ -866,11 +866,51 @@ function locateTranscript(cwd) {
|
|
|
866
866
|
if (entries.length === 0) {
|
|
867
867
|
throw new Error(`No transcripts in ${dir}. Pass --transcript <path> to pick one manually.`);
|
|
868
868
|
}
|
|
869
|
-
|
|
869
|
+
let best = null;
|
|
870
|
+
for (const name of entries) {
|
|
870
871
|
const full = path.join(dir, name);
|
|
871
|
-
|
|
872
|
-
|
|
873
|
-
|
|
872
|
+
const timestamp = lastImageTimestamp(full);
|
|
873
|
+
if (timestamp && (!best || timestamp > best.timestamp)) {
|
|
874
|
+
best = { full, timestamp };
|
|
875
|
+
}
|
|
876
|
+
}
|
|
877
|
+
if (!best) {
|
|
878
|
+
throw new Error(
|
|
879
|
+
`No pasted images found in any transcript under ${dir}. The user may not have pasted any, or the transcript format changed; ask for a file path instead.`
|
|
880
|
+
);
|
|
881
|
+
}
|
|
882
|
+
return best.full;
|
|
883
|
+
}
|
|
884
|
+
function lastImageTimestamp(transcriptPath) {
|
|
885
|
+
let raw;
|
|
886
|
+
try {
|
|
887
|
+
raw = fs.readFileSync(transcriptPath, "utf-8");
|
|
888
|
+
} catch {
|
|
889
|
+
return null;
|
|
890
|
+
}
|
|
891
|
+
let latest = null;
|
|
892
|
+
for (const line of raw.split("\n")) {
|
|
893
|
+
if (!line.includes('"image"')) {
|
|
894
|
+
continue;
|
|
895
|
+
}
|
|
896
|
+
let parsed;
|
|
897
|
+
try {
|
|
898
|
+
parsed = JSON.parse(line);
|
|
899
|
+
} catch {
|
|
900
|
+
continue;
|
|
901
|
+
}
|
|
902
|
+
const entry = parsed;
|
|
903
|
+
if (entry.message?.role !== "user" || !Array.isArray(entry.message.content)) {
|
|
904
|
+
continue;
|
|
905
|
+
}
|
|
906
|
+
const hasImage = entry.message.content.some(
|
|
907
|
+
(block) => block?.type === "image" && block.source?.type === "base64"
|
|
908
|
+
);
|
|
909
|
+
if (hasImage && entry.timestamp && (!latest || entry.timestamp > latest)) {
|
|
910
|
+
latest = entry.timestamp;
|
|
911
|
+
}
|
|
912
|
+
}
|
|
913
|
+
return latest;
|
|
874
914
|
}
|
|
875
915
|
function extractUserImages(transcriptPath) {
|
|
876
916
|
let raw;
|
|
@@ -926,7 +966,7 @@ function recoverPastedImages(options = {}) {
|
|
|
926
966
|
return { transcript, images };
|
|
927
967
|
}
|
|
928
968
|
const program = new Command();
|
|
929
|
-
program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("2.3.
|
|
969
|
+
program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("2.3.2");
|
|
930
970
|
program.command("analyze", { isDefault: true }).description("Analyze an image into structured JSON evidence (default command)").requiredOption("-i, --input <path|url>", "Input image path or https URL").option("-o, --output <path>", "Write result JSON to a file").option("-m, --model <name>", "Provider model name").option("-p, --provider <name>", `Vision provider (${listProviders().join(", ")})`).option("--prompt <text>", "Extra focus for this image").option("--timeout <ms>", "Provider timeout in milliseconds", "180000").option("--provider-bin <path>", "Provider binary path (default: agy)").option("--workdir <path>", "Working directory for the provider").action(async (options) => {
|
|
931
971
|
try {
|
|
932
972
|
const timeoutMs = Number.parseInt(options.timeout, 10);
|
package/package.json
CHANGED
package/skills/modlens/SKILL.md
CHANGED
|
@@ -66,7 +66,8 @@ Harnesses rarely hand you a clean path. Look for these signals:
|
|
|
66
66
|
- Codex wraps every pasted or attached image in a text tag like
|
|
67
67
|
`<image name=[Image #1] path="/tmp/xxxx.png">`. Extract the `path` value and run modlens on it. Pasted images live in a temp file the harness already created.
|
|
68
68
|
- A placeholder like `image content omitted because you do not support image input` means the harness stripped an image for you. The path tag next to it still holds the real file. Use it.
|
|
69
|
-
- Claude Code never writes pasted images to a regular temp file. Behind a text-only gateway you will only see a placeholder like `[Unsupported Image]` or `[Image #1]`, with no path. When that happens, run `modlens recover-paste` (add `--count <n>` for several images): it
|
|
69
|
+
- Claude Code never writes pasted images to a regular temp file. Behind a text-only gateway you will only see a placeholder like `[Unsupported Image]` or `[Image #1]`, with no path. When that happens, run `modlens recover-paste` (add `--count <n>` for several images): it recovers the pasted image bytes from the local session transcripts and prints real file paths as JSON. Feed that path to `modlens -i`.
|
|
70
|
+
You do not need to know which session you are; you cannot know it (no env var carries a session id) and the command does not take one. It scans every transcript of the current project directory and picks the session holding the newest pasted-image message by timestamp, which is necessarily the one the user just pasted into, even with concurrent sessions. Run it from the project directory the conversation is happening in. If recovery fails (transcript format is Claude Code internals and may change), fall back to asking the user to drag the image file into the terminal or type its path.
|
|
70
71
|
- If the user mentions an image but no tag or path appears anywhere in the message, ask for the file path instead of guessing.
|
|
71
72
|
|
|
72
73
|
## Workflow
|