@liustack/modlens 3.16.4 → 3.16.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,8 +1,18 @@
1
1
  # Changelog
2
2
 
3
+ ## 3.16.6 - 2026-08-15
4
+
5
+ - **A `null` where the contract asks for nothing no longer fails the read ([#37](https://github.com/liustack/modlens/issues/37)).** The reported failure named `visual.notes`, which the validator only reports when the field is present in a shape it does not accept: leaving it out was always fine. `null` is the shape a model reaches for when it has nothing to say, and it is what the reporter's own workaround had to legalize, so that is what this fixes. An optional field holding `null` is now dropped before the check rather than passing through it, which keeps the read alive and keeps `null` out of the fields this contract declares: each one is absent or holds its declared type, as the schema always promised. A key a gateway adds on its own is dropped the same way when it is null. On a required field `null` is still a violation. The error also stopped calling everything missing, since reading `missing: visual.notes` about a field that was right there sends you looking in the wrong place.
6
+ - **The openai provider can ask the gateway to enforce the contract ([#37](https://github.com/liustack/modlens/issues/37)).** `modlens config set openai.structuredOutput true` sends it as `response_format: json_schema` in the strict form those endpoints require: every property required, `additionalProperties: false`, and the ones this contract leaves optional made nullable. It is derived from the same schema the runtime checks against, so there is no second copy to keep in step, which is the part that made the reporter's own workaround expensive: they had to hand-write the whole thing to get thinking-disabled qwen through. Off by default, since a gateway without structured-output support answers 400 for the field, and a `response_format` set in `extraBody` still wins.
7
+
8
+ ## 3.16.5 - 2026-08-15
9
+
10
+ - **dsh: pasting into a plain text-only model works again ([#36](https://github.com/liustack/modlens/issues/36)).** The paste takeover shipped in 3.14.0 has been dead in every default install since 3.16.0 moved the verdict server-side. The verdict refuses when any model matching the selector label declares image input, which is right for a real vision model and wrong for the one case it could not see: this plugin's own `(modlens vision)` wrapper reuses the upstream model id verbatim and declares image input, because that declaration is exactly how the wrapper unlocks admission. So selecting plain `DeepSeek-V4-Pro` matched the real text-only model and the plugin's synthetic twin of it, the twin vetoed, and the paste fell through to dsh's own gate and its `MODEL_DOES_NOT_SUPPORT_IMAGES`. The verdict now skips a twin two ways, both requiring proof rather than a name: the provider ids this instance actually registered, tracked as each wrapper lands, and a model that carries the `(modlens vision)` marker on a provider id minted by the rule this plugin uses, which is how a sibling instance in the same process is recognized. A real vision provider still vetoes even if it borrows the marker, and an id someone else already holds is never trusted as ours. The verdict had no test at all, which is how a regression this total shipped and stayed for six releases; it now has an integration suite driving the real route against a registry shaped like a live install. Thanks to @Taz-dingo for a report that arrived with the conflicting rules already quoted side by side.
11
+ - **The docs cannot print a stale install command any more.** Every file the repo tracks is scanned, and any install of this package that is not pinned to an exact version has to carry `--config.minimumReleaseAge=0` as an argument of that same command. Review found the first version of that check passing five different ways of writing an unpinned install, including a command split across lines and a spec like `3.16.4+local` that only looks pinned.
12
+
3
13
  ## 3.16.4 - 2026-08-15
4
14
 
5
- - **How to update is written down, and the explanation it replaces was wrong.** 3.10.0 claimed that pnpm's release-age gate has a 10-day window and that an explicit version or dist-tag skips it, so every install command in this repo carried `@latest` as the fix. Measured on a machine with no gate configured at all: pnpm 11 turns `minimumReleaseAge` on by default at 24 hours (`24 * 60` in its config reader; `pnpm config get` does not surface that particular default), and `add @liustack/modlens@latest` installed 3.9.1 while 3.16.3 was the published latest. The gate filters the candidate versions before the tag is resolved, so the tag lands on an older one, and a day of held-back wall-clock time can be several releases on a fast week. The pnpm issue behind the original claim describes a bug in pnpm 10.16.1 that was fixed, and says in its own text that `@latest` was subject to the gate. What does work is naming the version, a deliberate request rather than a resolution: pnpm 11 installs it and records that one version as an approved exception, leaving everything else behind the window. Where a stricter policy is configured it refuses instead, and `--config.minimumReleaseAge=0` lifts the gate for one command, for every package that command resolves. `harness-setup` now has an Updating section covering both install shapes, including why `update` cannot cross a major (it stays inside the recorded semver range) and how to check what actually landed.
15
+ - **How to update is written down, and the explanation it replaces was wrong.** 3.10.0 claimed that pnpm's release-age gate has a 10-day window and that an explicit version or dist-tag skips it, so every install command in this repo carried `@latest` as the fix. Measured on a machine with no gate configured at all: pnpm 11 turns `minimumReleaseAge` on by default at 24 hours (`24 * 60` in its config reader; `pnpm config get` does not surface that particular default), and asking for `@latest` installed 3.9.1 while 3.16.3 was the published latest. The gate filters the candidate versions before the tag is resolved, so the tag lands on an older one, and a day of held-back wall-clock time can be several releases on a fast week. The pnpm issue behind the original claim describes a bug in pnpm 10.16.1 that was fixed, and says in its own text that `@latest` was subject to the gate. What does work is naming the version, a deliberate request rather than a resolution: pnpm 11 installs it and records that one version as an approved exception, leaving everything else behind the window. Where a stricter policy is configured it refuses instead, and `--config.minimumReleaseAge=0` lifts the gate for one command, for every package that command resolves. `harness-setup` now has an Updating section covering both install shapes, including why `update` cannot cross a major (it stays inside the recorded semver range) and how to check what actually landed.
6
16
 
7
17
  ## 3.16.3 - 2026-08-15
8
18
 
@@ -74,7 +84,7 @@
74
84
 
75
85
  ## 3.9.0 - 2026-08-13
76
86
 
77
- - **The first plug-in vision plugin for DeepSeek Harness (dsh).** The npm package is now also a dsh bundle: `dsh plugin --profile <name> add @liustack/modlens` is the whole install. It registers a native `read_image` tool (schema in every model request, so there is no trigger heuristic at all) that spawns the modlens CLI shipped in the same package, declares the vision schema as its canonical output contract, and renders evidence text for the model. Phase 2 rides `agent/pre-step`: images pasted or dropped into the dsh Web UI are read automatically and enter the step as modlens evidence blocks, with failed reads degrading to an explanatory note instead of rejecting the step (`autoRead: false` in the plugin row turns this off). The plugin imports no dsh packages (raw JSON-Schema tool registration, node builtins only), which is also the smallest possible surface against developer-preview churn. Verified end to end on a real dsh headless profile: the DeepSeek model called `read_image` and quoted the exact transcription back.
87
+ - **The first plug-in vision plugin for DeepSeek Harness (dsh).** The npm package is now also a dsh bundle, so a single `dsh plugin --profile <name> add` is the whole install (the current command, with its version named, is in [INSTALL.md](INSTALL.md)). It registers a native `read_image` tool (schema in every model request, so there is no trigger heuristic at all) that spawns the modlens CLI shipped in the same package, declares the vision schema as its canonical output contract, and renders evidence text for the model. Phase 2 rides `agent/pre-step`: images pasted or dropped into the dsh Web UI are read automatically and enter the step as modlens evidence blocks, with failed reads degrading to an explanatory note instead of rejecting the step (`autoRead: false` in the plugin row turns this off). The plugin imports no dsh packages (raw JSON-Schema tool registration, node builtins only), which is also the smallest possible surface against developer-preview churn. Verified end to end on a real dsh headless profile: the DeepSeek model called `read_image` and quoted the exact transcription back.
78
88
  - **Grok Build joins as the fifth reusable harness.** `reuse.grok` grants the local Grok CLI login as an engine: discovery reads `~/.grok` (OAuth evidence in auth.json, model ids from models_cache.json judged by the builtin vision table), and the route drives headless `grok -p` with `--json-schema` (which accepts this project's schema unmodified; the structuredOutput field carries the conforming answer) and `--allow Read`, following the claude-cli template since headless grok has no image-attach flag. Verified live: an exact OCR read through a real SuperGrok login. The agent region order becomes antigravity, codex, opencode, grok, pi-cli, claude-cli.
79
89
 
80
90
  ## 3.8.0 - 2026-08-13
package/README.md CHANGED
@@ -34,7 +34,7 @@ Issues are welcome any time: [open one](https://github.com/liustack/modlens/issu
34
34
 
35
35
  ## Highlights
36
36
 
37
- **🥇 The first vision plugin for DeepSeek Harness (dsh):** one command, `npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@latest`, and the text-only DeepSeek model behind dsh reads images through a native `modlens_read_image` tool. Updating is the same command again, with one wrinkle worth knowing: pnpm 11 holds back releases published in the last 24 hours, and `@latest` does not skip that, so a same-day fix needs its version named ([how to update](docs/harness-setup.md#keeping-it-up-to-date)).
37
+ **🥇 The first vision plugin for DeepSeek Harness (dsh):** one command, `npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.16.6`, and the text-only DeepSeek model behind dsh reads images through a native `modlens_read_image` tool. Updating is the same command again. The version is named rather than `@latest` on purpose: pnpm 11 holds back releases published in the last 24 hours and resolves the tag against what survives, so `@latest` would install whatever shipped a day ago ([details](docs/harness-setup.md#keeping-it-up-to-date)).
38
38
 
39
39
  Pasting an image works two ways. **① Just paste.** On a text-only model the pasted image lands as a private temp file and its path enters the composer — the same interaction OpenCode and Pi ship — and the `modlens_read_image` tool takes it from there. **② Pick a `(modlens vision)` entry** in the model selector (it remembers your choice, so once is enough), then paste: the thumbnail stays visible in your message, closer to the Codex app feel, and the image is converted to structured evidence at request time, answered by the same underlying route. The plugin auto-discovers every provider route carrying text-only DeepSeek or GLM models and adds a wrapped entry per route (a stock install gets **`DeepSeek-V4-Flash (modlens vision)`** and **`DeepSeek-V4-Pro (modlens vision)`**; extra routes like opencode-go or zai get their own); the two families' own vision models are excluded automatically. Which paste route applies is the host's per-model call: only a model its metadata positively confirms text-only is taken over, anything unconfirmed is left alone, so vision models keep their native paste ([details](docs/harness-setup.md)).
40
40
 
package/README.zh-CN.md CHANGED
@@ -34,7 +34,7 @@ DeepSeek 和 GLM 的主力对话模型是纯文本的,无法进行图片识别
34
34
 
35
35
  ## 亮点
36
36
 
37
- **🥇 全网第一个支持 DeepSeek Harness(dsh)的外挂视觉识别插件:**一条命令 `npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@latest`,dsh 背后的纯文本 DeepSeek 模型即可通过原生 `modlens_read_image` 工具读图。更新就是再跑一遍同一条命令,有一个坑值得知道:pnpm 11 会扣住最近 24 小时内发布的版本,而 `@latest` 跳不过去,所以想要当天的修复得点名版本号([怎么更新](docs/harness-setup.zh-CN.md#保持更新))。
37
+ **🥇 全网第一个支持 DeepSeek Harness(dsh)的外挂视觉识别插件:**一条命令 `npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.16.6`,dsh 背后的纯文本 DeepSeek 模型即可通过原生 `modlens_read_image` 工具读图。更新就是再跑一遍同一条命令。这里点名版本号而不用 `@latest` 是有意的:pnpm 11 会扣住最近 24 小时内发布的版本,dist-tag 只在剩下的里面解析,用 `@latest` 装到的会是一天前发布的那个([细节](docs/harness-setup.zh-CN.md#保持更新))。
38
38
 
39
39
  DeepSeek Harness 粘贴识图有两种玩法。
40
40
 
@@ -69,7 +69,7 @@ agy # 浏览器完成
69
69
  **DeepSeek Harness(dsh)用户不走 skill 流程**,本包就是原生 dsh 插件:
70
70
 
71
71
  ```sh
72
- npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@latest
72
+ npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.16.6
73
73
  ```
74
74
 
75
75
  装完即有 `modlens_read_image` 工具,选「(modlens vision)」模型变体即可直接粘贴识图。引擎配置同样在 `~/.modlens`,详见[宿主接入](docs/harness-setup.zh-CN.md)。
package/dist/main.js CHANGED
@@ -619,12 +619,71 @@ const VISION_RESULT_SCHEMA = {
619
619
  },
620
620
  required: ["summary", "ocr", "layout", "semantics", "visual", "uncertainty"]
621
621
  };
622
+ function strictSchema(node) {
623
+ if (node.type === "object") {
624
+ const properties = {};
625
+ const required = node.required ?? [];
626
+ for (const [key, child] of Object.entries(node.properties ?? {})) {
627
+ const strict = strictSchema(child);
628
+ properties[key] = required.includes(key) ? strict : (
629
+ // Strict mode has no optional properties, only nullable ones.
630
+ { anyOf: [strict, { type: "null" }] }
631
+ );
632
+ }
633
+ return {
634
+ type: "object",
635
+ properties,
636
+ required: Object.keys(properties),
637
+ additionalProperties: false
638
+ };
639
+ }
640
+ if (node.type === "array" && node.items) {
641
+ return { ...node, items: strictSchema(node.items) };
642
+ }
643
+ return node;
644
+ }
645
+ function visionResponseFormat() {
646
+ return {
647
+ type: "json_schema",
648
+ json_schema: {
649
+ name: "vision_result",
650
+ strict: true,
651
+ schema: strictSchema(VISION_RESULT_SCHEMA)
652
+ }
653
+ };
654
+ }
622
655
  function visionResultSchemaJson() {
623
656
  return JSON.stringify(VISION_RESULT_SCHEMA);
624
657
  }
625
658
  function missingSchemaFields(result) {
626
659
  return schemaViolations(VISION_RESULT_SCHEMA, result, "");
627
660
  }
661
+ function withoutEmptyOptionals(value, schema) {
662
+ if (schema.type === "object") {
663
+ if (typeof value !== "object" || value === null || Array.isArray(value)) {
664
+ return value;
665
+ }
666
+ const record = value;
667
+ const cleaned = {};
668
+ for (const [key, entry] of Object.entries(record)) {
669
+ const childSchema = schema.properties?.[key];
670
+ const isRequired = schema.required?.includes(key) ?? false;
671
+ if (entry === null && !isRequired) {
672
+ continue;
673
+ }
674
+ cleaned[key] = childSchema ? withoutEmptyOptionals(entry, childSchema) : entry;
675
+ }
676
+ return cleaned;
677
+ }
678
+ if (schema.type === "array" && schema.items && Array.isArray(value)) {
679
+ const itemSchema = schema.items;
680
+ return value.map((item) => withoutEmptyOptionals(item, itemSchema));
681
+ }
682
+ return value;
683
+ }
684
+ function normalizeVisionResult(result) {
685
+ return withoutEmptyOptionals(result, VISION_RESULT_SCHEMA);
686
+ }
628
687
  function schemaViolations(schema, value, path2) {
629
688
  const label = path2 || "(root)";
630
689
  if (schema.type === "object") {
@@ -1256,6 +1315,13 @@ ${JSON_TEMPLATE_INSTRUCTION}`;
1256
1315
  mergeExtraBody(
1257
1316
  {
1258
1317
  model,
1318
+ // Asked for, never assumed: a gateway without
1319
+ // structured-output support answers 400 for a field
1320
+ // it does not know (issue #37). A response_format the
1321
+ // caller supplied wins outright rather than being
1322
+ // merged into ours, since the two describe the same
1323
+ // thing and a blend of them describes neither.
1324
+ ...options.settings?.structuredOutput && options.settings?.extraBody?.response_format === void 0 ? { response_format: visionResponseFormat() } : {},
1259
1325
  messages: [
1260
1326
  {
1261
1327
  role: "user",
@@ -1286,14 +1352,15 @@ ${JSON_TEMPLATE_INSTRUCTION}`;
1286
1352
  if (!text) {
1287
1353
  throw new Error("OpenAI-compatible API returned no message content.");
1288
1354
  }
1289
- const result = extractJson(text);
1290
- if (result === null) {
1355
+ const rawResult = extractJson(text);
1356
+ if (rawResult === null) {
1291
1357
  throw new Error(`OpenAI-compatible API returned non-JSON output: ${truncate(text)}`);
1292
1358
  }
1359
+ const result = normalizeVisionResult(rawResult);
1293
1360
  const missing = missingSchemaFields(result);
1294
1361
  if (missing.length > 0) {
1295
1362
  throw new Error(
1296
- `OpenAI-compatible API returned JSON that does not match the vision schema (missing: ${missing.join(", ")}). Retry, or switch to gemini-api / anthropic for enforced schemas. Got: ${truncate(text)}`
1363
+ `OpenAI-compatible API returned JSON that does not match the vision schema (wrong or missing: ${missing.join(", ")}). Retry, or switch to gemini-api / anthropic for enforced schemas. Got: ${truncate(text)}`
1297
1364
  );
1298
1365
  }
1299
1366
  return {
@@ -1432,12 +1499,31 @@ function setConfigValue(dottedKey, value, configPath = CONFIG_PATH) {
1432
1499
  const dot = dottedKey.indexOf(".");
1433
1500
  if (dot <= 0 || dot === dottedKey.length - 1) {
1434
1501
  throw new Error(
1435
- `Invalid config key: ${dottedKey}. Use "provider", "reuse.<claude|codex|opencode|pi|grok>", "guards.<denyModels|allowModels|denyWhenUnknown>", or "<provider>.<apiKey|baseUrl|model|extraBody>".`
1502
+ `Invalid config key: ${dottedKey}. Use "provider", "proxy", "reuse.<claude|codex|opencode|pi|grok>", "guards.<denyModels|allowModels|denyWhenUnknown>", or "<provider>.<apiKey|baseUrl|model|proxy|extraBody|structuredOutput>".`
1436
1503
  );
1437
1504
  }
1438
1505
  const providerName = dottedKey.slice(0, dot);
1439
1506
  const field = dottedKey.slice(dot + 1);
1440
- if (field === "extraBody") {
1507
+ if (field === "structuredOutput") {
1508
+ if ((providerAliases()[providerName] ?? providerName) !== "openai") {
1509
+ throw new Error(
1510
+ `structuredOutput applies to the openai provider only, not ${providerName}.`
1511
+ );
1512
+ }
1513
+ const normalized = value.trim().toLowerCase();
1514
+ if (normalized !== "" && normalized !== "true" && normalized !== "false") {
1515
+ throw new Error(
1516
+ `${providerName}.structuredOutput must be true or false (empty clears).`
1517
+ );
1518
+ }
1519
+ config2.providers ??= {};
1520
+ config2.providers[providerName] ??= {};
1521
+ if (normalized === "") {
1522
+ delete config2.providers[providerName].structuredOutput;
1523
+ } else {
1524
+ config2.providers[providerName].structuredOutput = normalized === "true";
1525
+ }
1526
+ } else if (field === "extraBody") {
1441
1527
  config2.providers ??= {};
1442
1528
  config2.providers[providerName] ??= {};
1443
1529
  if (value.trim() === "") {
@@ -1450,7 +1536,7 @@ function setConfigValue(dottedKey, value, configPath = CONFIG_PATH) {
1450
1536
  }
1451
1537
  } else if (!STRING_FIELDS.includes(field)) {
1452
1538
  throw new Error(
1453
- `Unknown config field: ${field}. Use apiKey, baseUrl, model, proxy, or extraBody.`
1539
+ `Unknown config field: ${field}. Use apiKey, baseUrl, model, proxy, extraBody, or structuredOutput.`
1454
1540
  );
1455
1541
  } else {
1456
1542
  config2.providers ??= {};
@@ -1539,6 +1625,9 @@ function renderEffectiveConfig(config2, env = process.env) {
1539
1625
  fields[field] = `${shown} (${source})`;
1540
1626
  }
1541
1627
  }
1628
+ if (fileSettings.structuredOutput !== void 0) {
1629
+ fields.structuredOutput = `${fileSettings.structuredOutput} (file)`;
1630
+ }
1542
1631
  if (fileSettings.extraBody !== void 0) {
1543
1632
  fields.extraBody = `${JSON.stringify(fileSettings.extraBody)} (file)`;
1544
1633
  }
@@ -2770,10 +2859,11 @@ async function runProvider(provider, model, options, resolvedInput, timeoutMs, c
2770
2859
  `Provider ${provider.name} implements neither execute nor buildInvocation.`
2771
2860
  );
2772
2861
  }
2862
+ parsed.result = normalizeVisionResult(parsed.result);
2773
2863
  const missing = missingSchemaFields(parsed.result);
2774
2864
  if (missing.length > 0) {
2775
2865
  throw new Error(
2776
- `${provider.name} returned a result that does not match the vision schema (missing: ${missing.join(", ")}).`
2866
+ `${provider.name} returned a result that does not match the vision schema (wrong or missing: ${missing.join(", ")}).`
2777
2867
  );
2778
2868
  }
2779
2869
  return parsed;
@@ -4069,7 +4159,7 @@ function parsePositiveInt(raw, flag) {
4069
4159
  }
4070
4160
  return Number.parseInt(raw, 10);
4071
4161
  }
4072
- program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("3.16.4");
4162
+ program.name("modlens").description("Plug-in vision for text-only LLMs: image in, structured JSON evidence out").version("3.16.6");
4073
4163
  program.command("analyze", { isDefault: true }).description("Analyze an image into structured JSON evidence (default command)").requiredOption("-i, --input <path|url>", "Input image path or https URL").option("-o, --output <path>", "Write result JSON to a file").option("-m, --model <name>", "Provider model name").option("-p, --provider <name>", `Vision provider (${listProviders().join(", ")})`).option("--prompt <text>", "Extra focus for this image").option("--timeout <ms>", "Provider timeout in milliseconds", "180000").option("--provider-bin <path>", "Provider binary path (default: agy)").option("--workdir <path>", "Working directory for the provider").option(
4074
4164
  "--extra-body <json>",
4075
4165
  `JSON merged into the API request body, e.g. '{"thinking":{"type":"disabled"}}'`
@@ -4179,7 +4269,7 @@ program.command("doctor").description(
4179
4269
  configPath: CONFIG_PATH,
4180
4270
  // Lets doctor name an installed skill copy that is older than
4181
4271
  // the CLI reporting on it (issue #33).
4182
- version: "3.16.4"
4272
+ version: "3.16.6"
4183
4273
  });
4184
4274
  const output = options.json ? JSON.stringify(report, null, 2) : renderDoctorReport(report);
4185
4275
  process.stdout.write(`${output}
@@ -4199,9 +4289,9 @@ config.command("init").description(`Create a starter config at ${CONFIG_PATH}`).
4199
4289
  process.stdout.write(
4200
4290
  [
4201
4291
  `Created ${CONFIG_PATH}`,
4202
- "Everything is optional. Two things you can set:",
4292
+ "Everything is optional. The usual ones:",
4203
4293
  " modlens config set provider <name> which provider analyzes images",
4204
- " modlens config set <provider>.<apiKey|baseUrl|model> <value> provider credentials",
4294
+ " modlens config set <provider>.<apiKey|baseUrl|model> <value> provider settings",
4205
4295
  ` modlens config set <provider>.extraBody '{"thinking":{"type":"disabled"}}' vendor request fields`,
4206
4296
  ""
4207
4297
  ].join("\n")
package/docs/cli.md CHANGED
@@ -97,6 +97,6 @@ Five providers: `antigravity-cli` (no key), `gemini-api` (fastest free route), `
97
97
  Other subcommands:
98
98
 
99
99
  - `modlens guard [--model <id>]`: should the engine run for the active model at all? Exit 0 allow, 1 deny, verdict as JSON.
100
- - `modlens config <init|set|show>`: keys are `provider`, `proxy` (HTTP/HTTPS proxy for the API providers, `HTTPS_PROXY`/`HTTP_PROXY` also honored), `reuse.<claude|codex|opencode|pi|grok>`, `guards.<denyModels|allowModels|denyWhenUnknown>`, and `<provider>.<apiKey|baseUrl|model|proxy|extraBody>`.
100
+ - `modlens config <init|set|show>`: keys are `provider`, `proxy` (HTTP/HTTPS proxy for the API providers, `HTTPS_PROXY`/`HTTP_PROXY` also honored), `reuse.<claude|codex|opencode|pi|grok>`, `guards.<denyModels|allowModels|denyWhenUnknown>`, and `<provider>.<apiKey|baseUrl|model|proxy|extraBody>`, plus `openai.structuredOutput` (that route only).
101
101
  - `modlens doctor`: Node and node:sqlite, provider readiness, the failover chains for this machine, the detected harness, the guard's rules with a live verdict, and the Reuse section with per-harness grant decisions and discovered vision. Spends no quota; `--json` for a machine-readable report.
102
102
 
package/docs/cli.zh-CN.md CHANGED
@@ -94,5 +94,5 @@ modlens recover-paste # pull a pasted image into a fil
94
94
  其他子命令:
95
95
 
96
96
  - `modlens guard [--model <id>]`:判断当前激活的模型到底该不该运行引擎。退出码 0 表示放行,1 表示拒绝,判定结果以 JSON 输出。
97
- - `modlens config <init|set|show>`:可用的键有 `provider`、`proxy`(API provider 的 HTTP/HTTPS 代理,也认 `HTTPS_PROXY`/`HTTP_PROXY`)、`reuse.<claude|codex|opencode|pi|grok>`、`guards.<denyModels|allowModels|denyWhenUnknown>`,以及 `<provider>.<apiKey|baseUrl|model|proxy|extraBody>`。
97
+ - `modlens config <init|set|show>`:可用的键有 `provider`、`proxy`(API provider 的 HTTP/HTTPS 代理,也认 `HTTPS_PROXY`/`HTTP_PROXY`)、`reuse.<claude|codex|opencode|pi|grok>`、`guards.<denyModels|allowModels|denyWhenUnknown>`,以及 `<provider>.<apiKey|baseUrl|model|proxy|extraBody>`,另有 `openai.structuredOutput`(仅这条路线用得上)。
98
98
  - `modlens doctor`:报告 Node 与 node:sqlite、各 provider 的就绪状态、本机的故障转移链、检测到的 harness、guard 规则和一次现场判定,以及 Reuse 一节里按 harness 的授权决定与发现的视觉能力。不花任何额度,`--json` 输出机器可读报告。
@@ -55,7 +55,7 @@ OpenCode with DeepSeek: `opencode auth login`, pick DeepSeek and paste the key (
55
55
  dsh is different from the other harnesses: modlens plugs in as a native tool, not a prompt-triggered skill. The package itself is a dsh bundle, so one command installs it into a profile:
56
56
 
57
57
  ```sh
58
- npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@latest
58
+ npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.16.6
59
59
  ```
60
60
 
61
61
  This registers a `modlens_read_image` tool whose schema reaches the model on every request (no trigger heuristics), runs the modlens CLI shipped inside the same package, and returns the structured evidence as the tool's canonical JSON output. Engines, reuse grants, and guard rules stay in `~/.modlens/config.json`, shared with every other harness. dsh is in developer preview and its plugin surface may change; the plugin keeps its touch small (raw tool registration, the llm adapter surface for the vision variants, the attachment reader, and one agent pre-step hook) and degrades loudly if any of them moves.
@@ -63,35 +63,35 @@ This registers a `modlens_read_image` tool whose schema reaches the model on eve
63
63
  ### Keeping it up to date
64
64
 
65
65
  modlens ships often, and both install shapes freeze at whatever version they
66
- got. On dsh, re-run the install. `add` is the command, not `update`: `update`
67
- stays inside the semver range already recorded, and the range a plain install
68
- writes is a caret one, so a profile that once landed on 2.7.1 updates to 2.8.0
69
- and never crosses into 3.x.
66
+ got. On dsh, re-run the install with the version named:
70
67
 
71
68
  ```sh
72
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@latest
69
+ npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.6
73
70
  ```
74
71
 
75
- Restart dsh, then check what actually landed with
76
- `npx -y @deepseek-ai/dsh plugin --profile <name> list`. That check is worth
77
- running, because what you get may not be the newest release: pnpm 11 holds back
78
- anything published in the last 24 hours (`minimumReleaseAge`, on by default) and
79
- resolves silently to the newest release older than that. `@latest` does not
80
- change it: the gate filters the candidates before the tag is resolved, so the
81
- tag lands on an older one. What that costs is a day of wall-clock time, not one
82
- version, so on a fast-moving week it can be several releases back.
72
+ `npm view @liustack/modlens version` prints the current one, and this page is
73
+ stamped with it at release time.
83
74
 
84
- When it is not, name the version instead of asking for a tag:
75
+ Two things in that command are deliberate. `add`, not `update`, because
76
+ `update` stays inside the semver range already recorded, and a plain install
77
+ records a caret range, so a profile that once landed on 2.7.1 updates to 2.8.0
78
+ and never crosses into 3.x. And a named version rather than `@latest`, because
79
+ pnpm 11 holds back anything published in the last 24 hours
80
+ (`minimumReleaseAge`, on by default) and resolves the tag against what survives
81
+ that filter: `@latest` lands on an older release instead of skipping the gate.
82
+ What it costs is a day of wall-clock time rather than one version, which on a
83
+ fast-moving week is several releases. A named version is a deliberate request,
84
+ so pnpm installs it, and since 11.1.3 records that one version as an approved
85
+ exception in the profile's `pnpm-workspace.yaml`, leaving everything else
86
+ behind the window.
87
+
88
+ Restart dsh, then confirm what actually landed:
85
89
 
86
90
  ```sh
87
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.4
91
+ npx -y @deepseek-ai/dsh plugin --profile <name> list
88
92
  ```
89
93
 
90
- `npm view @liustack/modlens version` prints the current one. A named version is
91
- a deliberate request rather than a resolution, so pnpm 11 installs it, and
92
- since 11.1.3 records that one version as an approved exception in the profile's
93
- `pnpm-workspace.yaml`; everything else stays behind the window. The
94
- [troubleshooting page](troubleshooting.md#dsh-says-declares-no-dshbundle--installed-as-a-plain-dependency)
94
+ The [troubleshooting page](troubleshooting.md#dsh-says-declares-no-dshbundle--installed-as-a-plain-dependency)
95
95
  covers the stricter case, where you configured `minimumReleaseAge` yourself and
96
96
  pnpm refuses rather than approves.
97
97
 
@@ -55,28 +55,30 @@ OpenCode 接 DeepSeek:执行 `opencode auth login`,选择 DeepSeek 并粘贴
55
55
  dsh 与其他 harness 不同:modlens 以原生工具的形式接入,而不是靠提示词触发的 skill。本包自身就是一个 dsh bundle,一条命令即可装进某个 profile:
56
56
 
57
57
  ```sh
58
- npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@latest
58
+ npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.16.6
59
59
  ```
60
60
 
61
61
  这会注册一个 `modlens_read_image` 工具,它的 schema 随每次请求抵达模型(不靠触发启发式),运行同一个包里自带的 modlens CLI,并把结构化证据作为工具的标准 JSON 输出返回。引擎、复用授权和 guard 规则仍在 `~/.modlens/config.json` 里,与其他所有 harness 共享。dsh 还在开发者预览阶段,插件接口可能变化。这个插件刻意保持很小的接触面(原生工具注册、视觉变体所用的 llm 适配层、附件读取器,以及一个 agent 执行前钩子),其中任何一处变动,它都会大声报错而不是无声退化。
62
62
 
63
63
  ### 保持更新
64
64
 
65
- modlens 发布很频繁,而两种安装形态都会冻结在装进来的那个版本上。dsh 上重跑一遍安装即可。命令是 `add` 而不是 `update`:`update` 只在 package.json 里已记录的 semver 范围内挪动,而普通安装写进去的是 caret 范围,所以一个当初装到 2.7.1 的 profile 只会更新到 2.8.0,永远进不了 3.x。
65
+ modlens 发布很频繁,而两种安装形态都会冻结在装进来的那个版本上。dsh 上重跑一遍安装即可,版本号要点名:
66
66
 
67
67
  ```sh
68
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@latest
68
+ npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.6
69
69
  ```
70
70
 
71
- 重启 dsh,然后用 `npx -y @deepseek-ai/dsh plugin --profile <name> list` 看看实际装到了什么。这一步值得做,因为拿到的未必是最新版:pnpm 11 会扣住最近 24 小时内发布的版本(`minimumReleaseAge`,默认开启),静默解析到比它更早的那个最新版本。`@latest` 改变不了这一点:冷静期先把候选版本过滤掉,dist-tag 才在剩下的里面解析,于是它落到了更旧的那个上。代价是一天的自然时间,不是一个版本,发布密集的一周里可能落后好几个版本。
71
+ `npm view @liustack/modlens version` 可以查到当前版本号,本页的版本号则由发布流程自动写入。
72
72
 
73
- 需要今天的版本时,改成点名版本号,而不是要一个 tag:
73
+ 这条命令里有两处是刻意的。用 `add` 而不是 `update`,因为 `update` 只在已记录的 semver 范围内挪动,而普通安装写进去的是 caret 范围,所以一个当初装到 2.7.1 的 profile 只会更新到 2.8.0,永远进不了 3.x。点名版本号而不用 `@latest`,则是因为 pnpm 11 会扣住最近 24 小时内发布的版本(`minimumReleaseAge`,默认开启),dist-tag 只在通过过滤的候选里解析:`@latest` 会落到更旧的版本上,而不是跳过冷静期。代价是一天的自然时间,不是一个版本,发布密集的一周里就是好几个版本。点名版本是一次明确的指定,所以 pnpm 会装上它,11.1.3 起还会把这一个版本作为已批准的例外写进该 profile 的 `pnpm-workspace.yaml`,其余一切仍留在窗口后面。
74
+
75
+ 重启 dsh,然后确认实际装到了什么:
74
76
 
75
77
  ```sh
76
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.4
78
+ npx -y @deepseek-ai/dsh plugin --profile <name> list
77
79
  ```
78
80
 
79
- `npm view @liustack/modlens version` 可以查到当前版本号。点名版本是一次明确指定而不是一次解析,所以 pnpm 11 会装上它,11.1.3 起还会把这一个版本作为已批准的例外写进该 profile 的 `pnpm-workspace.yaml`,其余一切仍留在窗口后面。更严格的情况(你自己配过 `minimumReleaseAge`,pnpm 会拒绝而不是批准)见[故障排查](troubleshooting.zh-CN.md#dsh-提示-declares-no-dshbundle--installed-as-a-plain-dependency)。
81
+ 更严格的情况(你自己配过 `minimumReleaseAge`,pnpm 会拒绝而不是批准)见[故障排查](troubleshooting.zh-CN.md#dsh-提示-declares-no-dshbundle--installed-as-a-plain-dependency)。
80
82
 
81
83
  skill 类 harness 上,skill 是一个拷贝出来的文件夹,拷贝会保留安装时的版本,重跑安装原地覆盖即可。`modlens doctor` 会读出它能找到的每一份拷贝里钉住的版本,并标出落后于当前 CLI 的那些,让版本漂移在坑到人之前就先暴露出来。
82
84
 
@@ -71,6 +71,8 @@ The CLI prints one JSON object to stdout:
71
71
 
72
72
  Required fields: `summary`, `ocr`, `layout`, `semantics`, `visual`, `uncertainty` — every top-level field, `visual` included. (Earlier docs called `visual` optional; the enforced schema has always required it, so build to the schema.)
73
73
 
74
+ Optional fields: `ocr.lines[].language`, `semantics.intent`, `semantics.entities[].evidence`, `semantics.relations`, `visual.dominant_colors`, `visual.style`, `visual.notes`. Each is either absent or holds its declared type. Never `null`: a model with nothing to say there often writes one, and modlens drops the key before the result reaches you, so reading an optional field means checking whether it is there, not whether it is null.
75
+
74
76
  `layout.regions[].type` is a free string, not a closed list. Region kinds are an open set: a fixed enum rejected `link` on any web screenshot and `search` on a portal, and a rejected result fails the whole read over a descriptive label. The field's schema `description` names the common vocabulary as guidance, which reaches every provider that enforces this schema server-side, so an unlisted kind costs nothing.
75
77
 
76
78
  Changes from v1: pixel `bbox` coordinates and numeric `confidence` scores were removed. Vision models fabricate both, so v2 stops pretending to provide them.
@@ -71,6 +71,8 @@ CLI 向 stdout 打印一个 JSON 对象:
71
71
 
72
72
  必填字段:`summary`、`ocr`、`layout`、`semantics`、`visual`、`uncertainty`,也就是每一个顶层字段,`visual` 也不例外。(早期文档把 `visual` 写成可选,但强制执行的 schema 一直要求它,请以 schema 为准。)
73
73
 
74
+ 可选字段:`ocr.lines[].language`、`semantics.intent`、`semantics.entities[].evidence`、`semantics.relations`、`visual.dominant_colors`、`visual.style`、`visual.notes`。每一个要么不存在,要么就是它声明的类型,绝不会是 `null`:模型在这些位置没话可说时经常写 `null`,modlens 会在结果交到你手上之前把这个键删掉,所以读可选字段只需判断它在不在,不用判断是不是 null。
75
+
74
76
  `layout.regions[].type` 是自由字符串,不是封闭列表。区域类型本质是开放集合:固定枚举会让任何网页截图里的 `link`、门户页里的 `search` 直接落选,而一次落选就为了一个描述性标签废掉整次识别。常用词表写在该字段的 schema `description` 里作为指引,凡是在服务端强制执行这份 schema 的 provider 都会收到,没列到的类型不会有任何代价。
75
77
 
76
78
  相对 v1 的变化:删掉了像素级 `bbox` 坐标和数值型 `confidence` 分数。视觉模型会凭空编造这两样,v2 不再假装提供。
@@ -98,10 +98,36 @@ Working as intended. Codex writes pasted images to disk and puts the path in the
98
98
 
99
99
  ```
100
100
  OpenAI-compatible API returned JSON that does not match the vision schema
101
- (missing: ocr, ocr.full_text, ...)
101
+ (wrong or missing: visual.notes, ...)
102
102
  ```
103
103
 
104
- That endpoint returned a partial result. Only agy, gemini-api, anthropic, and claude-cli enforce the schema server-side, so weaker gateways can produce half a result. Retry once, then switch:
104
+ That endpoint returned something the contract does not accept. Note the
105
+ wording: a field named here can be absent, or present with the wrong shape.
106
+ For an optional field like `visual.notes`, only the second is possible, since
107
+ leaving it out is accepted. A `null` there is dropped rather than refused, so
108
+ what remains is a genuinely wrong type.
109
+
110
+ Most OpenAI-compatible gateways enforce nothing server-side, so the contract
111
+ travels as a filled-in JSON template in the prompt and a weaker model can
112
+ answer with half of it, especially with thinking turned off. Ask the gateway
113
+ to enforce it instead:
114
+
115
+ ```bash
116
+ modlens config set openai.structuredOutput true
117
+ ```
118
+
119
+ That sends the contract as `response_format: json_schema` in strict form,
120
+ derived from the same schema modlens checks against, so there is nothing to
121
+ keep in sync by hand. It is off by default because a gateway that does not
122
+ support the field answers 400 for it. If that happens:
123
+
124
+ ```bash
125
+ modlens config set openai.structuredOutput false
126
+ ```
127
+
128
+ A `response_format` you set yourself in `extraBody` wins over the derived one.
129
+
130
+ Failing that, retry once, then switch:
105
131
 
106
132
  ```bash
107
133
  modlens -i <image> -p gemini-api
@@ -137,7 +163,7 @@ simply lands on an older one. Name the exact version instead, which pnpm treats
137
163
  as a deliberate request rather than a resolution:
138
164
 
139
165
  ```sh
140
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.4
166
+ npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.6
141
167
  ```
142
168
 
143
169
  `npm view @liustack/modlens version` prints the current one. pnpm 11 installs a named
@@ -152,7 +178,7 @@ file:
152
178
 
153
179
  ```yaml
154
180
  minimumReleaseAgeExclude:
155
- - '@liustack/modlens@3.16.4'
181
+ - '@liustack/modlens@3.16.6'
156
182
  ```
157
183
 
158
184
  Or lift the gate for a single command, which lifts it for everything that
@@ -98,10 +98,26 @@ tag in the message carries its path.
98
98
 
99
99
  ```
100
100
  OpenAI-compatible API returned JSON that does not match the vision schema
101
- (missing: ocr, ocr.full_text, ...)
101
+ (wrong or missing: visual.notes, ...)
102
102
  ```
103
103
 
104
- 那个端点返回了残缺的结果。只有 agy、gemini-api、anthropic 和 claude-cli 在服务端强制执行 schema,较弱的网关可能只产出半个结果。重试一次,然后换 provider:
104
+ 那个端点返回了不符合契约的内容。注意措辞:被点名的字段可能是缺失,也可能是存在但形状不对。像 `visual.notes` 这样的可选字段只可能是后者,因为它缺失是被接受的。写成 `null` 也会被丢弃而不是拒绝,所以剩下的就是真正的类型错误。
105
+
106
+ 大多数 OpenAI 兼容网关在服务端什么都不强制,契约是以填好的 JSON 模板形式随提示词发过去的,能力弱一些的模型可能只答出一半,关掉思考时尤其明显。可以改成让网关自己强制执行:
107
+
108
+ ```bash
109
+ modlens config set openai.structuredOutput true
110
+ ```
111
+
112
+ 这会把契约以 `response_format: json_schema` 的严格形式发过去,schema 由 modlens 校验用的那份推导而来,没有需要手工同步的副本。默认关闭,因为不支持这个字段的网关会直接 400。真遇到就关回去:
113
+
114
+ ```bash
115
+ modlens config set openai.structuredOutput false
116
+ ```
117
+
118
+ 你自己在 `extraBody` 里设的 `response_format` 优先级更高。
119
+
120
+ 还是不行就重试一次,然后换 provider:
105
121
 
106
122
  ```bash
107
123
  modlens -i <image> -p gemini-api
@@ -128,7 +144,7 @@ dsh profile 装到的是旧版 modlens。`dsh.bundle` 声明从 3.9.0 起才存
128
144
  `@latest` 绕不开这一层,本页早先的说法是错的。冷静期先把候选版本过滤掉,dist-tag 才在剩下的里面解析,于是它直接落到了更旧的那个上。改成写死精确版本号,pnpm 会把它当作一次明确的指定,而不是一次解析:
129
145
 
130
146
  ```sh
131
- npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.4
147
+ npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.6
132
148
  ```
133
149
 
134
150
  `npm view @liustack/modlens version` 可以查到当前版本号。pnpm 11 会装上被点名的版本,11.1.3 起还会把它作为一条已批准的例外写进该 profile 的 `pnpm-workspace.yaml`,其余所有包和 modlens 以后的版本仍然留在窗口后面。
@@ -137,7 +153,7 @@ npx -y @deepseek-ai/dsh plugin --profile <name> add @liustack/modlens@3.16.4
137
153
 
138
154
  ```yaml
139
155
  minimumReleaseAgeExclude:
140
- - '@liustack/modlens@3.16.4'
156
+ - '@liustack/modlens@3.16.6'
141
157
  ```
142
158
 
143
159
  或者只为这一条命令解除冷静期,注意它解除的是这条命令解析到的所有包,不只 modlens:
package/dsh/index.js CHANGED
@@ -40,8 +40,15 @@ export function apply(ctx, config = {}) {
40
40
  if (config.autoRead === true) {
41
41
  registerAutoRead(ctx)
42
42
  }
43
+ // The provider ids this plugin registered itself. The takeover verdict has
44
+ // to skip them: our wrapper models are synthetic twins of upstream ones,
45
+ // carrying the upstream id and declaring image input, so a plain text-only
46
+ // label matches the twin and the twin's declaration vetoes the takeover
47
+ // that label deserved (issue #36). Filled by registerVisionProvider as
48
+ // wrappers land, including the later sweeps, and read by the verdict.
49
+ const ownProviders = new Set()
43
50
  if (config.visionProvider !== false) {
44
- registerVisionProvider(ctx, config)
51
+ registerVisionProvider(ctx, config, ownProviders)
45
52
  }
46
53
  // Paste-to-path: the browser half (dsh/client.js) intercepts image pastes
47
54
  // and POSTs the bytes here; the file lands in a private temp dir and the
@@ -57,7 +64,7 @@ export function apply(ctx, config = {}) {
57
64
  try {
58
65
  // scope carries webServer; the plugin's own ctx carries llm for the
59
66
  // takeover verdicts.
60
- registerPasteRoute(scope, ctx)
67
+ registerPasteRoute(scope, ctx, ownProviders)
61
68
  } catch (error) {
62
69
  console.error(`[modlens] paste-to-path route skipped: ${error}`)
63
70
  }
@@ -215,7 +222,14 @@ const PASTE_MAX_BYTES = 25 * 1024 * 1024
215
222
  * unresolvable answers false: the native path is the safe default, and a
216
223
  * text-only model merely keeps its old error message.
217
224
  */
218
- async function pasteTakeoverVerdict(host, label) {
225
+ // The provider ids registerVisionProvider mints: the legacy deepseek wrap and
226
+ // the `modlens-<upstream>` form auto-discovery uses. A sibling instance of
227
+ // this plugin derives its ids the same way, which is what makes the pair of
228
+ // checks below meaningful. A custom `config.providerId` is outside the
229
+ // convention on purpose and is covered by the registered-id set instead.
230
+ const OWN_PROVIDER_ID = /^(deepseek-modlens$|modlens-)/
231
+
232
+ async function pasteTakeoverVerdict(host, label, ownProviders) {
219
233
  if (typeof label !== 'string' || label.trim() === '') return false
220
234
  // Our own wrappers convert pastes at request time with the thumbnail
221
235
  // preserved; taking their paste over would defeat the better path.
@@ -229,6 +243,14 @@ async function pasteTakeoverVerdict(host, label) {
229
243
  for (const info of llm.listProviders()) {
230
244
  const providerId = info?.id
231
245
  if (!providerId) continue
246
+ // Our own wrapper: every model in it is a synthetic twin of an upstream
247
+ // one, carrying that upstream id and declaring image input because that
248
+ // is how the wrapper unlocks admission. Scanning it means a plain
249
+ // text-only label matches the twin by id and the twin vetoes the
250
+ // takeover the real model deserved (issue #36). Only ids this plugin
251
+ // registered itself are skipped, so a real vision provider, including
252
+ // one that happens to be named like ours, still votes.
253
+ if (ownProviders?.has(providerId)) continue
232
254
  let models = []
233
255
  try {
234
256
  models = await llm.listModels(providerId)
@@ -236,6 +258,21 @@ async function pasteTakeoverVerdict(host, label) {
236
258
  return false
237
259
  }
238
260
  for (const model of models) {
261
+ // A twin from another instance of this plugin, which the set above
262
+ // cannot know about: a second apply() in the same process hits the
263
+ // duplicate branch, does not claim the id, and would otherwise be
264
+ // vetoed by the first instance's wrapper. Both halves are required.
265
+ // The name marker alone proves nothing, since any provider can put that
266
+ // string in a model name and would then slip past a veto it deserves;
267
+ // the id is what makes it ours, because a sibling instance derives its
268
+ // provider id from the same rule this one does.
269
+ if (
270
+ OWN_PROVIDER_ID.test(providerId) &&
271
+ typeof model?.name === 'string' &&
272
+ /\(modlens vision\)/i.test(model.name)
273
+ ) {
274
+ continue
275
+ }
239
276
  for (const candidate of [model?.name, model?.id]) {
240
277
  if (typeof candidate !== 'string' || candidate.length === 0) continue
241
278
  if (!lowered.includes(candidate.toLowerCase())) continue
@@ -272,7 +309,7 @@ const PASTE_VERDICT_CAP = 32
272
309
  * stands down instead of swallowing pastes into a 404. Bound to the dsh web
273
310
  * server, which listens on loopback by default.
274
311
  */
275
- function registerPasteRoute(ctx, host) {
312
+ function registerPasteRoute(ctx, host, ownProviders) {
276
313
  const verdicts = new Map()
277
314
  // The cache key is only the selector label, which cannot tell two
278
315
  // same-named models on different routes apart. A route mounting mid-TTL
@@ -310,7 +347,7 @@ function registerPasteRoute(ctx, host) {
310
347
  let attempts = 0
311
348
  for (;;) {
312
349
  const startedEpoch = topologyEpoch
313
- takeover = await pasteTakeoverVerdict(host, label)
350
+ takeover = await pasteTakeoverVerdict(host, label, ownProviders)
314
351
  if (topologyEpoch === startedEpoch) {
315
352
  verdicts.delete(label)
316
353
  verdicts.set(label, { takeover, at: Date.now() })
@@ -398,7 +435,7 @@ function registerPasteRoute(ctx, host) {
398
435
  * deepseek-official wrap keeps its historical `deepseek-modlens` id, so a
399
436
  * selector remembering that provider survives the upgrade.
400
437
  */
401
- function registerVisionProvider(ctx, config) {
438
+ function registerVisionProvider(ctx, config, ownProviders) {
402
439
  // Wrap only the text-only members of these families. Their own vision
403
440
  // models (present or future: deepseek-vl/ocr/janus, glm-4.5v, glm-5v-...)
404
441
  // need no bridge and are excluded by name and by declared modality.
@@ -464,6 +501,11 @@ function registerVisionProvider(ctx, config) {
464
501
  },
465
502
  evidenceCache: new Map(),
466
503
  })
504
+ // Trusted as ours only on a registration this call actually made. A
505
+ // duplicate below means someone else holds that id, and skipping a
506
+ // provider we do not own would let a real vision model's paste be
507
+ // taken over, which is the bug the verdict exists to prevent.
508
+ ownProviders?.add(providerId)
467
509
  return true
468
510
  } catch (error) {
469
511
  // A duplicate means a concurrent or earlier registration already won:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@liustack/modlens",
3
- "version": "3.16.4",
3
+ "version": "3.16.6",
4
4
  "description": "Plug-in vision for text-only LLMs, powered by the free Antigravity CLI",
5
5
  "type": "module",
6
6
  "bin": {
@@ -20,11 +20,11 @@ powershell -ExecutionPolicy Bypass -File <skill-dir>\scripts\run.ps1 <args>
20
20
 
21
21
  It resolves a working runtime (PATH `modlens`, then `npx`, then `bunx`) and forwards your arguments unchanged. Exit 78 means no runtime: relay the `nextSteps` from its stderr JSON instead of retrying.
22
22
 
23
- If your harness forbids running scripts, reason through the same order by hand and run the first line that works (the pinned version is 3.16.4):
23
+ If your harness forbids running scripts, reason through the same order by hand and run the first line that works (the pinned version is 3.16.6):
24
24
 
25
- 1. A `modlens` on `PATH` whose major version is 3 and is at least 3.16.4: `modlens <args>`.
26
- 2. Otherwise, if `npx` exists: `npx --yes --package @liustack/modlens@3.16.4 modlens <args>`.
27
- 3. Otherwise, if `bunx` exists: `bunx --bun @liustack/modlens@3.16.4 <args>`.
25
+ 1. A `modlens` on `PATH` whose major version is 3 and is at least 3.16.6: `modlens <args>`.
26
+ 2. Otherwise, if `npx` exists: `npx --yes --package @liustack/modlens@3.16.6 modlens <args>`.
27
+ 3. Otherwise, if `bunx` exists: `bunx --bun @liustack/modlens@3.16.6 <args>`.
28
28
  4. Otherwise tell the user no JavaScript runtime was found and that installing Node 22.19+ (https://nodejs.org) or Bun (https://bun.sh) is the next step. Do not claim modlens itself failed.
29
29
 
30
30
  `references/runtime.md` documents the pin and the diagnostic fields.
@@ -12,14 +12,14 @@ Read this when the user asks how to set up, configure, or switch ModLens provide
12
12
  modlens config init # write a starter config (refuses to overwrite; --force to redo)
13
13
  modlens config show # effective file, API keys masked
14
14
  modlens config set provider <name> # change the default provider
15
- modlens config set <provider>.<field> <value> # fields: apiKey, baseUrl, model, extraBody
15
+ modlens config set <provider>.<field> <value> # fields: apiKey, baseUrl, model, proxy, extraBody, structuredOutput
16
16
  ```
17
17
 
18
18
  `config set` writes the file with 0600 permissions.
19
19
 
20
20
  ## The file's exact shape
21
21
 
22
- Everything lives under four top-level keys, all optional. This example shows every supported key and field at once (a real file only needs what you use). A missing file means all defaults. Provider settings sit under `providers.<name>`, not at the top level, which is the mistake hand-editors make most.
22
+ Everything lives under five top-level keys, all optional. This example shows every supported key and field at once (a real file only needs what you use). A missing file means all defaults. Provider settings sit under `providers.<name>`, not at the top level, which is the mistake hand-editors make most.
23
23
 
24
24
  ```json
25
25
  {
@@ -42,7 +42,9 @@ Everything lives under four top-level keys, all optional. This example shows eve
42
42
  "apiKey": "sk-...",
43
43
  "baseUrl": "https://dashscope.aliyuncs.com/compatible-mode/v1",
44
44
  "model": "qwen3.6-27b",
45
- "extraBody": { "thinking": { "type": "disabled" } }
45
+ "proxy": "http://127.0.0.1:7890",
46
+ "extraBody": { "thinking": { "type": "disabled" } },
47
+ "structuredOutput": true
46
48
  },
47
49
  "anthropic": {
48
50
  "apiKey": "sk-ant-...",
@@ -57,8 +59,9 @@ Everything lives under four top-level keys, all optional. This example shows eve
57
59
  Field semantics:
58
60
 
59
61
  - `provider`: which provider runs when `-p` is not given. Canonical names or aliases both work (`agy`/`antigravity` for `antigravity-cli`, `gemini` for `gemini-api`, `openai-compat` for `openai`, `claude` for `anthropic`, `claude-code` for `claude-cli`). Empty or absent pins nothing: the failover chain decides, trying configured API providers before the agent CLIs.
60
- - `providers.<name>.<field>`: four fields exist, `apiKey`, `baseUrl`, `model`, and `extraBody`. Every provider entry is optional, and every field inside it is optional. Alias keys are read too (settings saved under `gemini` are found when `gemini-api` resolves), with the canonical key winning on conflict.
61
- - `providers.<name>.extraBody`: a JSON object merged into the request body of the API providers (`gemini-api`, `openai`, `anthropic`), for whatever knobs that vendor has and modlens has no flag for. Turning thinking off is the usual reason, see the section below. Nested objects merge key by key, so adding one knob leaves the rest of that block alone. The fields carrying the image, the prompt, and the schema enforcement are refused with an error naming the field. The two CLI providers take no request body, so a run on `antigravity-cli` or `claude-cli` ignores it and says so in `meta.warnings`.
62
+ - `providers.<name>.<field>`: six fields exist, `apiKey`, `baseUrl`, `model`, `proxy`, `extraBody`, and `structuredOutput` (the openai route only). Every provider entry is optional, and every field inside it is optional. Alias keys are read too (settings saved under `gemini` are found when `gemini-api` resolves), with the canonical key winning on conflict.
63
+ - `providers.<name>.extraBody`: a JSON object merged into the request body of the API providers (`gemini-api`, `openai`, `anthropic`), for whatever knobs that vendor has and modlens has no flag for. Turning thinking off is the usual reason, see the section below. Nested objects merge key by key, so adding one knob leaves the rest of that block alone. The fields carrying the image, the prompt, and each route's own enforcement machinery are refused with an error naming the field. `response_format` on the `openai` route is not one of them: setting it there deliberately replaces the schema modlens would otherwise send. The two CLI providers take no request body, so a run on `antigravity-cli` or `claude-cli` ignores it and says so in `meta.warnings`.
64
+ - `providers.openai.structuredOutput`: `true` asks an OpenAI-compatible gateway to enforce the vision contract itself, as `response_format: json_schema` in the strict form those endpoints require. Off by default, since a gateway without structured-output support answers 400 for the field. A `response_format` you set in `extraBody` wins over it.
62
65
  - `guards`: the invocation guard, for people who run both text-only and vision-capable models through the same client. Both lists hold glob patterns (`*` and `?`, case-insensitive, matched against the model name and `provider/model`), set with `modlens config set guards.denyModels '["gemini-3*"]'` or `guards.allowModels` (a JSON array or a comma-separated list, empty clears). Two ways to express the same intent, pick the shorter list:
63
66
  - `denyModels` alone: everything runs the engine except the listed vision models. Right when text-only models are the majority of what you plug in.
64
67
  - `allowModels` non-empty (allowlist mode): only the listed models run the engine, every other identified model is denied. Right for the actual 2026 landscape, where text-only models are the short list. A deny pattern still wins over an allow match, so a broad allow can have its vision variants carved out, as in the example above: `glm-5.*` allows the text line while `glm-*v*` catches `glm-5v-turbo`. Anchor allow patterns tightly (`deepseek-v4-*`, not `deepseek*`) so a vendor's next multimodal generation falls off the list and steps aside until you have checked it.
@@ -105,7 +108,15 @@ modlens config set openai.apiKey <sk-key>
105
108
  modlens config set openai.model qwen3.6-27b
106
109
  ```
107
110
 
108
- For official OpenAI: baseUrl `https://api.openai.com/v1`, a vision-capable model. Environment equivalents: `OPENAI_BASE_URL`, `OPENAI_API_KEY`. The model must be multimodal; text-only models will fail or hallucinate. This route has no server-side schema enforcement, so occasional shape failures are surfaced as explicit errors; retry or switch provider.
111
+ For official OpenAI: baseUrl `https://api.openai.com/v1`, a vision-capable model. Environment equivalents: `OPENAI_BASE_URL`, `OPENAI_API_KEY`. The model must be multimodal; text-only models will fail or hallucinate.
112
+
113
+ This route enforces nothing server-side by default, so a weaker model can answer with half the contract and the run fails with an explicit error. If that happens, ask the gateway to enforce it:
114
+
115
+ ```bash
116
+ modlens config set openai.structuredOutput true
117
+ ```
118
+
119
+ The contract goes out as `response_format: json_schema` in strict form, derived from the schema modlens checks against. Off by default because a gateway without structured-output support answers 400 for the field, so turn it back off if the endpoint refuses it. Turning thinking off (below) makes the shape failures more likely, so the two often go together.
109
120
 
110
121
  ### anthropic (Claude API key)
111
122
 
@@ -12,14 +12,14 @@
12
12
  modlens config init # 写入一份起步配置(已存在则拒绝,--force 重写)
13
13
  modlens config show # 生效的配置文件,API key 打码显示
14
14
  modlens config set provider <name> # 更改默认 provider
15
- modlens config set <provider>.<field> <value> # 字段:apiKey、baseUrl、model、extraBody
15
+ modlens config set <provider>.<field> <value> # 字段:apiKey、baseUrl、model、proxy、extraBody、structuredOutput
16
16
  ```
17
17
 
18
18
  `config set` 写文件时权限为 0600。
19
19
 
20
20
  ## 配置文件的完整形状
21
21
 
22
- 所有内容都在四个顶层键之下,全部可选。下面的示例一次性展示了所有支持的键和字段(真实文件只需要写你用到的部分)。文件不存在就全用默认值。provider 的设置放在 `providers.<name>` 下面,不在顶层,手工编辑最常犯的就是这个错。
22
+ 所有内容都在五个顶层键之下,全部可选。下面的示例一次性展示了所有支持的键和字段(真实文件只需要写你用到的部分)。文件不存在就全用默认值。provider 的设置放在 `providers.<name>` 下面,不在顶层,手工编辑最常犯的就是这个错。
23
23
 
24
24
  ```json
25
25
  {
@@ -42,7 +42,9 @@ modlens config set <provider>.<field> <value> # 字段:apiKey、baseUrl、mo
42
42
  "apiKey": "sk-...",
43
43
  "baseUrl": "https://dashscope.aliyuncs.com/compatible-mode/v1",
44
44
  "model": "qwen3.6-27b",
45
- "extraBody": { "thinking": { "type": "disabled" } }
45
+ "proxy": "http://127.0.0.1:7890",
46
+ "extraBody": { "thinking": { "type": "disabled" } },
47
+ "structuredOutput": true
46
48
  },
47
49
  "anthropic": {
48
50
  "apiKey": "sk-ant-...",
@@ -57,8 +59,9 @@ modlens config set <provider>.<field> <value> # 字段:apiKey、baseUrl、mo
57
59
  字段含义:
58
60
 
59
61
  - `provider`:不传 `-p` 时由哪个 provider 执行。标准名和别名都行(`agy`/`antigravity` 对应 `antigravity-cli`,`gemini` 对应 `gemini-api`,`openai-compat` 对应 `openai`,`claude` 对应 `anthropic`,`claude-code` 对应 `claude-cli`)。留空或缺失表示不钉任何一个:由失败切换链决定,已配置的 API provider 先于 agent CLI 被尝试。
60
- - `providers.<name>.<field>`:共四个字段,`apiKey`、`baseUrl`、`model`、`extraBody`。每个 provider 条目都可选,条目里的每个字段也都可选。别名键同样会被读取(存在 `gemini` 下的设置在解析到 `gemini-api` 时也能找到),冲突时标准键胜出。
61
- - `providers.<name>.extraBody`:一个 JSON 对象,合并进 API provider(`gemini-api`、`openai`、`anthropic`)的请求体,用来传厂商有而 modlens 没有对应参数的开关。最常见的用途是关掉思考,见下文小节。嵌套对象逐键合并,所以加一个开关不会动到该块里的其他内容。承载图片、提示词和 schema 约束的字段会被拒绝,报错会点名该字段。两个 CLI provider 不发请求体,所以在 `antigravity-cli` 或 `claude-cli` 上运行时它会被忽略,并在 `meta.warnings` 里说明。
62
+ - `providers.<name>.<field>`:共六个字段,`apiKey`、`baseUrl`、`model`、`proxy`、`extraBody`、`structuredOutput`(仅 openai 路线)。每个 provider 条目都可选,条目里的每个字段也都可选。别名键同样会被读取(存在 `gemini` 下的设置在解析到 `gemini-api` 时也能找到),冲突时标准键胜出。
63
+ - `providers.<name>.extraBody`:一个 JSON 对象,合并进 API provider(`gemini-api`、`openai`、`anthropic`)的请求体,用来传厂商有而 modlens 没有对应参数的开关。最常见的用途是关掉思考,见下文小节。嵌套对象逐键合并,所以加一个开关不会动到该块里的其他内容。承载图片、提示词和各路线自身强制机制的字段会被拒绝,报错会点名该字段。`openai` 路线上的 `response_format` 不在此列:在那里设置它就是有意替换掉 modlens 本来会发的那份 schema。两个 CLI provider 不发请求体,所以在 `antigravity-cli` 或 `claude-cli` 上运行时它会被忽略,并在 `meta.warnings` 里说明。
64
+ - `providers.openai.structuredOutput`:设为 `true` 时,让 OpenAI 兼容网关自己强制执行视觉契约,以 `response_format: json_schema` 的严格形式发出。默认关闭,因为不支持结构化输出的网关会对这个字段返回 400。你在 `extraBody` 里设的 `response_format` 优先级更高。
62
65
  - `guards`:调用 guard,给在同一个客户端里既跑纯文本模型又跑视觉模型的人用。两个列表都放 glob 模式(支持 `*` 和 `?`,不区分大小写,同时匹配模型名和 `provider/model`),用 `modlens config set guards.denyModels '["gemini-3*"]'` 或 `guards.allowModels` 设置(JSON 数组或逗号分隔的列表都行,传空则清除)。两种写法表达同一个意图,选列表更短的那种:
63
66
  - 只用 `denyModels`:除了列出的视觉模型,其余全部运行引擎。适合你接入的模型大多是纯文本的情况。
64
67
  - `allowModels` 非空(白名单模式):只有列出的模型运行引擎,其他所有已识别的模型一律拒绝。适合 2026 年的实际格局,纯文本模型才是那份短名单。deny 模式仍然优先于 allow 匹配,所以宽泛的 allow 可以把视觉变体剔出去,正如上面的示例:`glm-5.*` 放行文本系列,`glm-*v*` 抓住 `glm-5v-turbo`。allow 模式要锚定得紧一些(写 `deepseek-v4-*` 而不是 `deepseek*`),这样厂商下一代多模态型号会自动掉出名单,等你检查过再上场。
@@ -105,7 +108,15 @@ modlens config set openai.apiKey <sk-key>
105
108
  modlens config set openai.model qwen3.6-27b
106
109
  ```
107
110
 
108
- 官方 OpenAI 的写法:baseUrl 用 `https://api.openai.com/v1`,配一个具备视觉能力的模型。对应的环境变量:`OPENAI_BASE_URL`、`OPENAI_API_KEY`。模型必须是多模态的,纯文本模型会失败或产生幻觉。这条路线没有服务端 schema 约束,偶发的结构错误会以明确报错的形式暴露出来,重试或换 provider 即可。
111
+ 官方 OpenAI 的写法:baseUrl 用 `https://api.openai.com/v1`,配一个具备视觉能力的模型。对应的环境变量:`OPENAI_BASE_URL`、`OPENAI_API_KEY`。模型必须是多模态的,纯文本模型会失败或产生幻觉。
112
+
113
+ 这条路线默认在服务端不做任何约束,能力弱一些的模型可能只答出契约的一半,运行就会以明确报错失败。真遇到就让网关自己强制执行:
114
+
115
+ ```bash
116
+ modlens config set openai.structuredOutput true
117
+ ```
118
+
119
+ 契约会以 `response_format: json_schema` 的严格形式发出去,schema 由 modlens 校验用的那份推导而来。默认关闭,因为不支持结构化输出的网关会对这个字段返回 400,端点拒绝就关回去。关掉思考(见下)会让结构错误更容易出现,所以这两项常常一起用。
109
120
 
110
121
  ### anthropic(Claude API key)
111
122
 
@@ -8,7 +8,7 @@ shell syntax.
8
8
 
9
9
  ## Pinned version
10
10
 
11
- - Pinned CLI version: 3.16.4
11
+ - Pinned CLI version: 3.16.6
12
12
  - npm package: `@liustack/modlens`
13
13
  - CLI binary name: `modlens`
14
14
 
@@ -24,7 +24,7 @@ $ErrorActionPreference = 'Stop'
24
24
  # package.json version, and the release script rewrites it on every bump.
25
25
  $Package = '@liustack/modlens'
26
26
  $Bin = 'modlens'
27
- $Pinned = '3.16.4'
27
+ $Pinned = '3.16.6'
28
28
  # -------------------------------------------------------------------------------
29
29
 
30
30
  $NativeNote = 'no native artifact is published for this tool yet; phase A ships npm launch paths only'
@@ -22,7 +22,7 @@ set -eu
22
22
  # package.json version, and the release script rewrites it on every bump.
23
23
  PKG="@liustack/modlens"
24
24
  BIN="modlens"
25
- PINNED="3.16.4"
25
+ PINNED="3.16.6"
26
26
  # -------------------------------------------------------------------------------
27
27
 
28
28
  NATIVE_NOTE="no native artifact is published for this tool yet; phase A ships npm launch paths only"