visual-ai-assertions 0.25.0 → 0.27.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +146 -76
- package/dist/index.cjs +225 -137
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +47 -13
- package/dist/index.d.ts +47 -13
- package/dist/index.js +225 -137
- package/dist/index.js.map +1 -1
- package/package.json +2 -5
package/README.md
CHANGED
|
@@ -2,6 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, Qwen, and GLM via OpenRouter — and get structured, typed results.
|
|
4
4
|
|
|
5
|
+
Every supported model is benchmarked against hand-labelled screenshots — **[see how they compare](https://nullp2ike.github.io/visual-reasoning/)** ([methodology](#benchmarks)).
|
|
6
|
+
|
|
5
7
|
## Installation
|
|
6
8
|
|
|
7
9
|
```bash
|
|
@@ -99,7 +101,7 @@ const ai = visualAI();
|
|
|
99
101
|
|
|
100
102
|
// Explicit configuration
|
|
101
103
|
const ai = visualAI({
|
|
102
|
-
model: "claude-sonnet-
|
|
104
|
+
model: "claude-sonnet-5-5", // optional, sensible defaults per provider
|
|
103
105
|
apiKey: "sk-...", // optional, defaults to provider env var
|
|
104
106
|
debug: true, // optional, logs prompts/responses to stderr
|
|
105
107
|
maxTokens: 4096, // optional, default 4096
|
|
@@ -110,7 +112,7 @@ const ai = visualAI({
|
|
|
110
112
|
|
|
111
113
|
// Use constants for IDE autocomplete
|
|
112
114
|
const ai = visualAI({
|
|
113
|
-
model: Model.Anthropic.
|
|
115
|
+
model: Model.Anthropic.SONNET_5_5,
|
|
114
116
|
});
|
|
115
117
|
```
|
|
116
118
|
|
|
@@ -198,8 +200,8 @@ import { writeFileSync } from "node:fs";
|
|
|
198
200
|
// Basic comparison
|
|
199
201
|
const result = await ai.compare(before, after);
|
|
200
202
|
|
|
201
|
-
// gemini-3-flash
|
|
202
|
-
// Pass { diffImage: false } to opt out.
|
|
203
|
+
// Every Gemini flash model — the gemini-3.8-flash default included — auto-includes
|
|
204
|
+
// an annotated diff. Pass { diffImage: false } to opt out.
|
|
203
205
|
|
|
204
206
|
// With custom prompt and instructions
|
|
205
207
|
const result = await ai.compare(before, after, {
|
|
@@ -207,8 +209,9 @@ const result = await ai.compare(before, after, {
|
|
|
207
209
|
instructions: ["Ignore date/time differences"],
|
|
208
210
|
});
|
|
209
211
|
|
|
210
|
-
//
|
|
211
|
-
//
|
|
212
|
+
// Requesting the diff image explicitly (it is already on by default for the flash tier).
|
|
213
|
+
// Supported on gemini-3-flash-preview, 3.5, 3.6, 3.7 and 3.8-flash (DIFF_ALLOWED_MODELS);
|
|
214
|
+
// Flash-Lite and Pro tiers produce none even when asked.
|
|
212
215
|
const result = await ai.compare(before, after, {
|
|
213
216
|
diffImage: true,
|
|
214
217
|
});
|
|
@@ -224,7 +227,7 @@ if (result.diffImage) {
|
|
|
224
227
|
pass: boolean; // true if no critical/major changes
|
|
225
228
|
reasoning: string; // overall summary
|
|
226
229
|
changes: ChangeEntry[]; // list of visual differences
|
|
227
|
-
diffImage?: { // present when diffing is enabled explicitly or by Gemini
|
|
230
|
+
diffImage?: { // present when diffing is enabled explicitly or by the Gemini flash default
|
|
228
231
|
data: Buffer; // PNG image data
|
|
229
232
|
width: number;
|
|
230
233
|
height: number;
|
|
@@ -513,15 +516,16 @@ The `VisualAIKnownError` union and `isVisualAIKnownError()` helper are useful wh
|
|
|
513
516
|
|
|
514
517
|
### Optional Configuration
|
|
515
518
|
|
|
516
|
-
| Variable | Description
|
|
517
|
-
| ---------------------------- |
|
|
518
|
-
| `VISUAL_AI_MODEL` |
|
|
519
|
-
| `
|
|
520
|
-
| `
|
|
521
|
-
| `
|
|
522
|
-
| `
|
|
523
|
-
| `
|
|
524
|
-
| `
|
|
519
|
+
| Variable | Description |
|
|
520
|
+
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
521
|
+
| `VISUAL_AI_MODEL` | Model used when `model` is not set in config, overriding the provider's default. It also **selects the provider**, since the provider is inferred from the model name — so this is how you pick a provider when several API keys are set. |
|
|
522
|
+
| `VISUAL_AI_REASONING_EFFORT` | Reasoning effort when `reasoningEffort` is not set in config. One of `minimal`, `low`, `medium`, `high`, `xhigh` (case-insensitive). An unrecognised value throws `VisualAIConfigError` rather than being ignored. |
|
|
523
|
+
| `VISUAL_AI_DEBUG` | Enable error diagnostic logging to stderr. Does **not** enable prompt/response logging. Use `"true"` or `"1"`. |
|
|
524
|
+
| `VISUAL_AI_DEBUG_PROMPT` | Enable prompt-only debug logging to stderr. Use `"true"` or `"1"`. |
|
|
525
|
+
| `VISUAL_AI_DEBUG_RESPONSE` | Enable response-only debug logging to stderr. Use `"true"` or `"1"`. |
|
|
526
|
+
| `VISUAL_AI_DEBUG_FRAMES` | Persist sampled video frames to disk for offline inspection. Use `"true"` or `"1"`. Frames are written to `./visual-ai-debug-frames/<timestamp>-<id>/` (override path with the next variable). Has no effect on image-only inputs. |
|
|
527
|
+
| `VISUAL_AI_DEBUG_FRAMES_DIR` | Override the base directory for `VISUAL_AI_DEBUG_FRAMES`. Each call still gets its own timestamped subdirectory inside it. |
|
|
528
|
+
| `VISUAL_AI_TRACK_USAGE` | Enable usage tracking (token counts and cost) to stderr. Use `"true"` or `"1"`. |
|
|
525
529
|
|
|
526
530
|
## Configuration
|
|
527
531
|
|
|
@@ -566,12 +570,12 @@ type SupportedMimeType = "image/jpeg" | "image/png" | "image/webp" | "image/gif"
|
|
|
566
570
|
|
|
567
571
|
**Default models:**
|
|
568
572
|
|
|
569
|
-
| Provider | Default Model
|
|
570
|
-
| ---------- |
|
|
571
|
-
| Anthropic | `claude-sonnet-
|
|
572
|
-
| OpenAI | `gpt-
|
|
573
|
-
| Google | `gemini-3-flash
|
|
574
|
-
| OpenRouter | `
|
|
573
|
+
| Provider | Default Model |
|
|
574
|
+
| ---------- | --------------------- |
|
|
575
|
+
| Anthropic | `claude-sonnet-5-5` |
|
|
576
|
+
| OpenAI | `gpt-6.1-sol` |
|
|
577
|
+
| Google | `gemini-3.8-flash` |
|
|
578
|
+
| OpenRouter | `meta/muse-spark-1.3` |
|
|
575
579
|
|
|
576
580
|
## Reasoning Effort
|
|
577
581
|
|
|
@@ -579,19 +583,22 @@ Control how deeply the model reasons before responding. Higher effort produces m
|
|
|
579
583
|
|
|
580
584
|
```typescript
|
|
581
585
|
const ai = visualAI({
|
|
582
|
-
reasoningEffort: "high", // "low" | "medium" | "high" | "xhigh"
|
|
586
|
+
reasoningEffort: "high", // "minimal" | "low" | "medium" | "high" | "xhigh"
|
|
583
587
|
});
|
|
584
588
|
```
|
|
585
589
|
|
|
586
|
-
|
|
590
|
+
Set `VISUAL_AI_REASONING_EFFORT` to apply one without touching code; an explicit `reasoningEffort` wins over it.
|
|
591
|
+
|
|
592
|
+
When omitted, each provider uses its default behavior. The `"xhigh"` level enables maximum reasoning depth. `"minimal"` is **not portable** — several OpenAI models reject it with HTTP 400 (including the `gpt-6.1-sol` default), and Google/OpenRouter clamp it to `low` rather than send it; use `"low"` as the floor unless you know the model accepts it.
|
|
587
593
|
|
|
588
|
-
| Provider
|
|
589
|
-
|
|
|
590
|
-
| Anthropic (Fable 5/Opus 4.8/4.7
|
|
591
|
-
| Anthropic (
|
|
592
|
-
|
|
|
593
|
-
|
|
|
594
|
-
|
|
|
594
|
+
| Provider | Native Parameter | `"xhigh"` maps to |
|
|
595
|
+
| -------------------------------------------------------------------- | ----------------------------------------------------- | -------------------- |
|
|
596
|
+
| Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5, Haiku 5.5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
|
|
597
|
+
| Anthropic (Opus 4.6, Sonnet 4.6) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
|
|
598
|
+
| Anthropic (Haiku 4.5) | budget-based extended thinking (token budget) | 16384-token budget |
|
|
599
|
+
| OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
|
|
600
|
+
| Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
|
|
601
|
+
| OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
|
|
595
602
|
|
|
596
603
|
## Supported Models
|
|
597
604
|
|
|
@@ -599,49 +606,60 @@ All listed models support image/vision input. Pass any model ID to the `model` c
|
|
|
599
606
|
|
|
600
607
|
### Anthropic
|
|
601
608
|
|
|
602
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
603
|
-
| ----------------- | ------------------- | ------------ | ------------- |
|
|
604
|
-
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work
|
|
605
|
-
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price
|
|
606
|
-
| Claude Opus
|
|
607
|
-
| Claude Opus
|
|
608
|
-
| Claude Opus 4.
|
|
609
|
-
| Claude
|
|
610
|
-
| Claude
|
|
611
|
-
| Claude
|
|
609
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
610
|
+
| ----------------- | ------------------- | ------------ | ------------- | ------------------------------------------------- |
|
|
611
|
+
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work |
|
|
612
|
+
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price |
|
|
613
|
+
| Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5 |
|
|
614
|
+
| Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh` |
|
|
615
|
+
| Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh` |
|
|
616
|
+
| Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier |
|
|
617
|
+
| Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output |
|
|
618
|
+
| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh` |
|
|
619
|
+
| Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work |
|
|
620
|
+
| Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation |
|
|
621
|
+
| Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | Newest Haiku; cheapest Claude; supports `xhigh` † |
|
|
622
|
+
| Claude Haiku 4.5 | `claude-haiku-4-5` | $1 | $5 | Fastest, budget-friendly |
|
|
623
|
+
|
|
624
|
+
† Haiku 5.5 prices apply to prompts up to 100K tokens; longer prompts bill at $0.50 / $2.50, far beyond screenshot-sized calls. Unlike Haiku 4.5 it uses adaptive thinking like the other current Claude models, and it thinks even when `reasoningEffort` is not set.
|
|
612
625
|
|
|
613
626
|
### OpenAI
|
|
614
627
|
|
|
615
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
616
|
-
| ------------- | --------------- | ------------ | ------------- |
|
|
617
|
-
| GPT-6 Astra | `gpt-6-astra` | $10 | $50 | Most capable; restricted access¹
|
|
618
|
-
| GPT-
|
|
619
|
-
| GPT-
|
|
620
|
-
| GPT-
|
|
621
|
-
| GPT-5.
|
|
622
|
-
| GPT-5.
|
|
623
|
-
| GPT-5.
|
|
624
|
-
| GPT-5.
|
|
625
|
-
| GPT-5.4
|
|
626
|
-
| GPT-5.4
|
|
627
|
-
| GPT-5
|
|
628
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
629
|
+
| ------------- | --------------- | ------------ | ------------- | ----------------------------------- |
|
|
630
|
+
| GPT-6 Astra | `gpt-6-astra` | $10 | $50 | Most capable; restricted access¹ |
|
|
631
|
+
| GPT-6.1 Sol | `gpt-6.1-sol` | $2 | $10 | **Default** — near-Astra quality² |
|
|
632
|
+
| GPT-6 Sol | `gpt-6-sol` | $2 | $10 | GPT-6 generation, frontier tier |
|
|
633
|
+
| GPT-6 Luna | `gpt-6-luna` | $0.10 | $0.50 | GPT-6 generation, fastest/cheapest |
|
|
634
|
+
| GPT-5.6 Sol | `gpt-5.6-sol` | $5 | $30 | Previous flagship, frontier tier |
|
|
635
|
+
| GPT-5.6 Terra | `gpt-5.6-terra` | $2 | $12 | Newest balanced, everyday tier |
|
|
636
|
+
| GPT-5.6 Luna | `gpt-5.6-luna` | $0.20 | $1.20 | Prior default — fast and cheap |
|
|
637
|
+
| GPT-5.5 | `gpt-5.5` | $5 | $30 | Previous flagship, 1M context |
|
|
638
|
+
| GPT-5.4 Pro | `gpt-5.4-pro` | $30 | $180 | Most capable, extended context |
|
|
639
|
+
| GPT-5.4 | `gpt-5.4` | $2.50 | $15 | Best vision quality |
|
|
640
|
+
| GPT-5.2 | `gpt-5.2` | $1.75 | $14 | Balanced quality and cost |
|
|
641
|
+
| GPT-5.4 mini | `gpt-5.4-mini` | $0.75 | $4.50 | Prior default — fast and affordable |
|
|
642
|
+
| GPT-5.4 nano | `gpt-5.4-nano` | $0.20 | $1.25 | Cheapest older-generation option |
|
|
643
|
+
| GPT-5 mini | `gpt-5-mini` | $0.25 | $2 | Fast and cheap |
|
|
628
644
|
|
|
629
645
|
¹ GPT-6 Astra is rolling out through OpenAI's Trusted Access Program, so many API keys cannot reach it yet — expect a `VisualAIProviderError` naming the model until your account is enabled.
|
|
630
646
|
|
|
631
|
-
Astra
|
|
647
|
+
Astra uses the same output budget as other OpenAI models: the 4096 default, raised to 16384 automatically at `high`/`xhigh`. It used to get 32768 at every effort because plain `ask()` calls exhausted the default, which looked like heavy reasoning. That was actually the image `ask()` schema bug fixed in this release: the model spent under 50 reasoning tokens, then printed whitespace until the budget ran out. With the fix, `ask()` and `check()` both completed every call at 4096 in live testing, and no call on the `golden` bench produced more than 548 output tokens. It also accepts a fifth reasoning level, `max`, above `xhigh`; this library's `reasoningEffort` stops at `xhigh`, which is passed through unchanged, so `max` is not currently reachable.
|
|
648
|
+
|
|
649
|
+
² GPT-6.1 Sol accepts `low`, `medium`, `high`, `xhigh` and `max`; `minimal` and `none` return HTTP 400, so use `low` as the floor.
|
|
632
650
|
|
|
633
651
|
### Google
|
|
634
652
|
|
|
635
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
636
|
-
| --------------------- | ------------------------ | ------------ | ------------- |
|
|
637
|
-
| Gemini 3.8 Flash | `gemini-3.8-flash` | $0.75 | $3.75 |
|
|
638
|
-
| Gemini 3.7 Flash | `gemini-3.7-flash` | $0.75 | $3.75 | Prior GA flash; intro pricing¹
|
|
639
|
-
| Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $7.50 | Prior GA flash; fewer out-tokens
|
|
640
|
-
| Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $9 | Strongest agentic & coding model
|
|
641
|
-
| Gemini 3.5 Flash Lite | `gemini-3.5-flash-lite` | $0.30 | $2.50 | GA — fast, cheap, agentic tier
|
|
642
|
-
| Gemini 3.1 Pro | `gemini-3.1-pro-preview` | $2 | $12 | Preview — most advanced reasoning
|
|
643
|
-
| Gemini 3.1 Flash Lite | `gemini-3.1-flash-lite` | $0.25 | $1.50 | GA — lightweight and cheap
|
|
644
|
-
| Gemini 3 Flash | `gemini-3-flash-preview` | $0.50 | $3 |
|
|
653
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
654
|
+
| --------------------- | ------------------------ | ------------ | ------------- | ---------------------------------- |
|
|
655
|
+
| Gemini 3.8 Flash | `gemini-3.8-flash` | $0.75 | $3.75 | **Default** — newest GA flash¹ |
|
|
656
|
+
| Gemini 3.7 Flash | `gemini-3.7-flash` | $0.75 | $3.75 | Prior GA flash; intro pricing¹ |
|
|
657
|
+
| Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $7.50 | Prior GA flash; fewer out-tokens |
|
|
658
|
+
| Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $9 | Strongest agentic & coding model |
|
|
659
|
+
| Gemini 3.5 Flash Lite | `gemini-3.5-flash-lite` | $0.30 | $2.50 | GA — fast, cheap, agentic tier |
|
|
660
|
+
| Gemini 3.1 Pro | `gemini-3.1-pro-preview` | $2 | $12 | Preview — most advanced reasoning |
|
|
661
|
+
| Gemini 3.1 Flash Lite | `gemini-3.1-flash-lite` | $0.25 | $1.50 | GA — lightweight and cheap |
|
|
662
|
+
| Gemini 3 Flash | `gemini-3-flash-preview` | $0.50 | $3 | Prior default; cheapest flash tier |
|
|
645
663
|
|
|
646
664
|
¹ Gemini 3.8 Flash and 3.7 Flash introductory pricing runs through 2026-12-31; both revert to $1.50 / $7.50 per MTok on 2027-01-01.
|
|
647
665
|
|
|
@@ -649,26 +667,78 @@ Astra reasons heavily enough to spend the entire 4096-token default output budge
|
|
|
649
667
|
|
|
650
668
|
Any [OpenRouter](https://openrouter.ai/models) model slug (always `vendor/model`) is accepted — the vendor prefix is how the library recognizes an OpenRouter model. The models below are tested and have pricing built in. Note that OpenRouter may route a request to different upstream hosts with different quantizations; keep that in mind when comparing benchmark numbers.
|
|
651
669
|
|
|
652
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
653
|
-
| -------------- | --------------------------- | ------------ | ------------- |
|
|
654
|
-
| Muse Spark 1.3 | `meta/muse-spark-1.3` | $1.25 | $4.25 | Meta flagship
|
|
655
|
-
| Grok 4.6 | `x-ai/grok-4.6` | $2 | $6 | Newest xAI flagship, 500K context
|
|
656
|
-
| Grok 4.5 | `x-ai/grok-4.5` | $2 | $6 | Prior xAI flagship, 500K context
|
|
657
|
-
| Kimi K3 | `moonshotai/kimi-k3` | $3 | $15 | Moonshot flagship, 1M context
|
|
658
|
-
| Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | $0.82 | $3.75 | Agentic/coding tier with vision
|
|
659
|
-
| Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input
|
|
660
|
-
| Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading
|
|
661
|
-
| Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 |
|
|
662
|
-
| GLM 5.3 Flash | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | Z.ai flash tier, 1.3M context²
|
|
670
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
671
|
+
| -------------- | --------------------------- | ------------ | ------------- | ----------------------------------- |
|
|
672
|
+
| Muse Spark 1.3 | `meta/muse-spark-1.3` | $1.25 | $4.25 | **Default** — Meta flagship; gated¹ |
|
|
673
|
+
| Grok 4.6 | `x-ai/grok-4.6` | $2 | $6 | Newest xAI flagship, 500K context |
|
|
674
|
+
| Grok 4.5 | `x-ai/grok-4.5` | $2 | $6 | Prior xAI flagship, 500K context |
|
|
675
|
+
| Kimi K3 | `moonshotai/kimi-k3` | $3 | $15 | Moonshot flagship, 1M context |
|
|
676
|
+
| Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | $0.82 | $3.75 | Agentic/coding tier with vision⁵ |
|
|
677
|
+
| Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input⁴ |
|
|
678
|
+
| Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading⁴ |
|
|
679
|
+
| Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 | Prior default — cheap flash vision |
|
|
680
|
+
| GLM 5.3 Flash | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | Z.ai flash tier, 1.3M context² |
|
|
681
|
+
| MiMo V2.6 Pro | `xiaomi/mimo-v2.6-pro` | $0.435 | $0.87 | Xiaomi flagship, 1M context³ |
|
|
663
682
|
|
|
664
683
|
¹ Muse Spark 1.3 is age-gated by OpenRouter: calls return HTTP 403 (`VisualAIAuthError`) until the account completes the 18+ confirmation at [openrouter.ai/settings/preferences](https://openrouter.ai/settings/preferences). It also reasons by default — expect several hundred reasoning tokens per call even with no `reasoningEffort` set.
|
|
665
684
|
|
|
666
685
|
² GLM 5.3 Flash reasons by default — expect one to two hundred reasoning tokens per call even with no `reasoningEffort` set, billed at the output rate. OpenRouter's own context cap for it is 1,048,576 tokens (Z.ai lists 1,310,720) and its output ceiling is 131,072.
|
|
667
686
|
|
|
687
|
+
³ MiMo V2.6 Pro also reasons by default: about 390 reasoning tokens per call with no `reasoningEffort` set, and about 790 at `medium`. OpenRouter serves it from two fp8 hosts at the same price, Xiaomi (~36 tok/s) and DeepInfra (~4 tok/s), so per-call latency varies widely with the host it is routed to.
|
|
688
|
+
|
|
689
|
+
⁴ Qwen3.8 Max and Qwen3.7 Plus reason past the 4096-token default on a large share of calls (4,500–5,400 reasoning tokens on the long ones), so **they get a 32768-token output budget automatically** at every reasoning effort. At the default, about half their `ask()` calls truncated in live testing; with the larger budget every call completed. Passing `maxTokens` explicitly still wins, and a call that used the whole budget would cost at most about $0.20 on Qwen3.8 Max.
|
|
690
|
+
|
|
691
|
+
⁵ Kimi K2.7 Code fails roughly one call in ten by answering in prose instead of JSON. At the 4096 default those calls surface as `VisualAITruncationError`, and with a larger `maxTokens` they finish and throw `VisualAIResponseParseError` instead, so raising the budget does not help. Retry failed calls.
|
|
692
|
+
|
|
668
693
|
Meta also publishes `meta/muse-spark-1.3-contributor`, the same model at $0.10 / $0.20 per MTok — about 12x cheaper — because Meta uses everything submitted through it for product improvement. It has **no named constant** (`Model.OpenRouter` does not expose it) and never appears by default anywhere in this library, so using it takes a deliberate, explicit choice: pass the slug directly as a plain string, `visualAI({ model: "meta/muse-spark-1.3-contributor" })`. Any OpenRouter slug works this way — see the note above the table — and cost tracking works correctly once you opt in. OpenRouter itself blocks it with HTTP 404 (`paid-model-training-violation-by-account`) until the account's privacy settings allow training endpoints, at [openrouter.ai/settings/privacy](https://openrouter.ai/settings/privacy). Only use it if sending your screenshots to Meta for training is a trade you've deliberately made.
|
|
669
694
|
|
|
670
695
|
`qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
|
|
671
696
|
|
|
697
|
+
## Benchmarks
|
|
698
|
+
|
|
699
|
+
Which model should you actually pick? This repo benchmarks every supported model
|
|
700
|
+
against hand-labelled screenshots and publishes the results:
|
|
701
|
+
|
|
702
|
+
**📊 [Live results — nullp2ike.github.io/visual-reasoning](https://nullp2ike.github.io/visual-reasoning/)**
|
|
703
|
+
|
|
704
|
+
The site carries the defect-discovery leaderboard, a screenshot × model matrix you
|
|
705
|
+
can drill into for any individual answer, and the same runs graded independently by
|
|
706
|
+
five different judges so you can see where the grading itself is contested.
|
|
707
|
+
|
|
708
|
+
The benchmark asks each model one open question — "What looks visually broken on
|
|
709
|
+
this page?" — with no hints, and an LLM judge matches what it reports against the
|
|
710
|
+
defects seeded into each screenshot. The headline is recall of those defects,
|
|
711
|
+
weighed against the extras a model invents per run: a missed defect is a bug that
|
|
712
|
+
ships, and an invented one is noise someone has to triage.
|
|
713
|
+
|
|
714
|
+
### The `golden` dataset
|
|
715
|
+
|
|
716
|
+
18 screenshots — 17 with exactly one deliberately seeded defect, plus a clean
|
|
717
|
+
control where anything reported counts as a false positive. The prompt names a
|
|
718
|
+
short list of out-of-scope non-defects (edge-clipped carousel items, content cut
|
|
719
|
+
off by the viewport bottom) so models are not penalised for reporting framing as
|
|
720
|
+
breakage.
|
|
721
|
+
|
|
722
|
+
Current coverage, at `medium` effort with 5 repeats per cell:
|
|
723
|
+
|
|
724
|
+
- 38 model/effort/fidelity series over 3,420 graded runs
|
|
725
|
+
|
|
726
|
+
A few results worth knowing before you choose a default:
|
|
727
|
+
|
|
728
|
+
| Model | Discovery recall | Extras / run | Cost / run |
|
|
729
|
+
| ----------------------------------------- | ---------------- | ------------ | ---------- |
|
|
730
|
+
| `claude-opus-5-5` | 100% | 0.53 | $0.023 |
|
|
731
|
+
| `claude-sonnet-5-5` _(Anthropic default)_ | 94% | 2.81 | $0.011 |
|
|
732
|
+
| `gpt-6.1-sol` _(OpenAI default)_ | 93% | 0.02 | $0.0055 |
|
|
733
|
+
| `gemini-3.8-flash` _(Google default)_ | 87% | 0.34 | $0.0050 |
|
|
734
|
+
|
|
735
|
+
Recall is not the whole picture. `claude-sonnet-5-5` matches the Fable tier on
|
|
736
|
+
recall but reports 2.81 extras per run against `gpt-6.1-sol`'s 0.02 — for
|
|
737
|
+
open-ended discovery that is a lot of noise to triage.
|
|
738
|
+
|
|
739
|
+
Full methodology, how to run a sweep, and how to add your own dataset:
|
|
740
|
+
[`bench/README.md`](bench/README.md).
|
|
741
|
+
|
|
672
742
|
## License
|
|
673
743
|
|
|
674
744
|
MIT
|