visual-ai-assertions 0.26.0 → 0.27.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +71 -21
- package/dist/index.cjs +9 -1
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +1 -0
- package/dist/index.d.ts +1 -0
- package/dist/index.js +9 -1
- package/dist/index.js.map +1 -1
- package/package.json +1 -3
package/README.md
CHANGED
|
@@ -2,6 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, Qwen, and GLM via OpenRouter — and get structured, typed results.
|
|
4
4
|
|
|
5
|
+
Every supported model is benchmarked against hand-labelled screenshots — **[see how they compare](https://nullp2ike.github.io/visual-reasoning/)** ([methodology](#benchmarks)).
|
|
6
|
+
|
|
5
7
|
## Installation
|
|
6
8
|
|
|
7
9
|
```bash
|
|
@@ -589,14 +591,14 @@ Set `VISUAL_AI_REASONING_EFFORT` to apply one without touching code; an explicit
|
|
|
589
591
|
|
|
590
592
|
When omitted, each provider uses its default behavior. The `"xhigh"` level enables maximum reasoning depth. `"minimal"` is **not portable** — several OpenAI models reject it with HTTP 400 (including the `gpt-6.1-sol` default), and Google/OpenRouter clamp it to `low` rather than send it; use `"low"` as the floor unless you know the model accepts it.
|
|
591
593
|
|
|
592
|
-
| Provider
|
|
593
|
-
|
|
|
594
|
-
| Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
|
|
595
|
-
| Anthropic (Opus 4.6, Sonnet 4.6)
|
|
596
|
-
| Anthropic (Haiku 4.5)
|
|
597
|
-
| OpenAI
|
|
598
|
-
| Google
|
|
599
|
-
| OpenRouter
|
|
594
|
+
| Provider | Native Parameter | `"xhigh"` maps to |
|
|
595
|
+
| -------------------------------------------------------------------- | ----------------------------------------------------- | -------------------- |
|
|
596
|
+
| Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5, Haiku 5.5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
|
|
597
|
+
| Anthropic (Opus 4.6, Sonnet 4.6) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
|
|
598
|
+
| Anthropic (Haiku 4.5) | budget-based extended thinking (token budget) | 16384-token budget |
|
|
599
|
+
| OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
|
|
600
|
+
| Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
|
|
601
|
+
| OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
|
|
600
602
|
|
|
601
603
|
## Supported Models
|
|
602
604
|
|
|
@@ -604,19 +606,22 @@ All listed models support image/vision input. Pass any model ID to the `model` c
|
|
|
604
606
|
|
|
605
607
|
### Anthropic
|
|
606
608
|
|
|
607
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
608
|
-
| ----------------- | ------------------- | ------------ | ------------- |
|
|
609
|
-
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work
|
|
610
|
-
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price
|
|
611
|
-
| Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5
|
|
612
|
-
| Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh`
|
|
613
|
-
| Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh`
|
|
614
|
-
| Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier
|
|
615
|
-
| Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output
|
|
616
|
-
| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh`
|
|
617
|
-
| Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work
|
|
618
|
-
| Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation
|
|
619
|
-
| Claude Haiku
|
|
609
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
610
|
+
| ----------------- | ------------------- | ------------ | ------------- | ------------------------------------------------- |
|
|
611
|
+
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work |
|
|
612
|
+
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price |
|
|
613
|
+
| Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5 |
|
|
614
|
+
| Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh` |
|
|
615
|
+
| Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh` |
|
|
616
|
+
| Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier |
|
|
617
|
+
| Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output |
|
|
618
|
+
| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh` |
|
|
619
|
+
| Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work |
|
|
620
|
+
| Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation |
|
|
621
|
+
| Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | Newest Haiku; cheapest Claude; supports `xhigh` † |
|
|
622
|
+
| Claude Haiku 4.5 | `claude-haiku-4-5` | $1 | $5 | Fastest, budget-friendly |
|
|
623
|
+
|
|
624
|
+
† Haiku 5.5 prices apply to prompts up to 100K tokens; longer prompts bill at $0.50 / $2.50, far beyond screenshot-sized calls. Unlike Haiku 4.5 it uses adaptive thinking like the other current Claude models, and it thinks even when `reasoningEffort` is not set.
|
|
620
625
|
|
|
621
626
|
### OpenAI
|
|
622
627
|
|
|
@@ -689,6 +694,51 @@ Meta also publishes `meta/muse-spark-1.3-contributor`, the same model at $0.10 /
|
|
|
689
694
|
|
|
690
695
|
`qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
|
|
691
696
|
|
|
697
|
+
## Benchmarks
|
|
698
|
+
|
|
699
|
+
Which model should you actually pick? This repo benchmarks every supported model
|
|
700
|
+
against hand-labelled screenshots and publishes the results:
|
|
701
|
+
|
|
702
|
+
**📊 [Live results — nullp2ike.github.io/visual-reasoning](https://nullp2ike.github.io/visual-reasoning/)**
|
|
703
|
+
|
|
704
|
+
The site carries the defect-discovery leaderboard, a screenshot × model matrix you
|
|
705
|
+
can drill into for any individual answer, and the same runs graded independently by
|
|
706
|
+
five different judges so you can see where the grading itself is contested.
|
|
707
|
+
|
|
708
|
+
The benchmark asks each model one open question — "What looks visually broken on
|
|
709
|
+
this page?" — with no hints, and an LLM judge matches what it reports against the
|
|
710
|
+
defects seeded into each screenshot. The headline is recall of those defects,
|
|
711
|
+
weighed against the extras a model invents per run: a missed defect is a bug that
|
|
712
|
+
ships, and an invented one is noise someone has to triage.
|
|
713
|
+
|
|
714
|
+
### The `golden` dataset
|
|
715
|
+
|
|
716
|
+
18 screenshots — 17 with exactly one deliberately seeded defect, plus a clean
|
|
717
|
+
control where anything reported counts as a false positive. The prompt names a
|
|
718
|
+
short list of out-of-scope non-defects (edge-clipped carousel items, content cut
|
|
719
|
+
off by the viewport bottom) so models are not penalised for reporting framing as
|
|
720
|
+
breakage.
|
|
721
|
+
|
|
722
|
+
Current coverage, at `medium` effort with 5 repeats per cell:
|
|
723
|
+
|
|
724
|
+
- 38 model/effort/fidelity series over 3,420 graded runs
|
|
725
|
+
|
|
726
|
+
A few results worth knowing before you choose a default:
|
|
727
|
+
|
|
728
|
+
| Model | Discovery recall | Extras / run | Cost / run |
|
|
729
|
+
| ----------------------------------------- | ---------------- | ------------ | ---------- |
|
|
730
|
+
| `claude-opus-5-5` | 100% | 0.53 | $0.023 |
|
|
731
|
+
| `claude-sonnet-5-5` _(Anthropic default)_ | 94% | 2.81 | $0.011 |
|
|
732
|
+
| `gpt-6.1-sol` _(OpenAI default)_ | 93% | 0.02 | $0.0055 |
|
|
733
|
+
| `gemini-3.8-flash` _(Google default)_ | 87% | 0.34 | $0.0050 |
|
|
734
|
+
|
|
735
|
+
Recall is not the whole picture. `claude-sonnet-5-5` matches the Fable tier on
|
|
736
|
+
recall but reports 2.81 extras per run against `gpt-6.1-sol`'s 0.02 — for
|
|
737
|
+
open-ended discovery that is a lot of noise to triage.
|
|
738
|
+
|
|
739
|
+
Full methodology, how to run a sweep, and how to add your own dataset:
|
|
740
|
+
[`bench/README.md`](bench/README.md).
|
|
741
|
+
|
|
692
742
|
## License
|
|
693
743
|
|
|
694
744
|
MIT
|
package/dist/index.cjs
CHANGED
|
@@ -497,6 +497,7 @@ var Model = {
|
|
|
497
497
|
SONNET_5_5: "claude-sonnet-5-5",
|
|
498
498
|
SONNET_5: "claude-sonnet-5",
|
|
499
499
|
SONNET_4_6: "claude-sonnet-4-6",
|
|
500
|
+
HAIKU_5_5: "claude-haiku-5-5",
|
|
500
501
|
HAIKU_4_5: "claude-haiku-4-5"
|
|
501
502
|
},
|
|
502
503
|
OpenAI: {
|
|
@@ -716,7 +717,8 @@ var XHIGH_CAPABLE_MODELS = /* @__PURE__ */ new Set([
|
|
|
716
717
|
Model.Anthropic.OPUS_4_8,
|
|
717
718
|
Model.Anthropic.OPUS_4_7,
|
|
718
719
|
Model.Anthropic.SONNET_5_5,
|
|
719
|
-
Model.Anthropic.SONNET_5
|
|
720
|
+
Model.Anthropic.SONNET_5,
|
|
721
|
+
Model.Anthropic.HAIKU_5_5
|
|
720
722
|
]);
|
|
721
723
|
function mapEffort(level, model) {
|
|
722
724
|
if (level === "minimal") return "low";
|
|
@@ -1428,6 +1430,12 @@ var PRICING_TABLE = {
|
|
|
1428
1430
|
inputPricePerToken: 3 / PER_MILLION,
|
|
1429
1431
|
outputPricePerToken: 15 / PER_MILLION
|
|
1430
1432
|
},
|
|
1433
|
+
// Prompts above 100K tokens bill at $0.50/$2.50, far beyond screenshot-sized
|
|
1434
|
+
// calls. Cached input is $0.01/MTok and cache writes $0.125/MTok (not modelled).
|
|
1435
|
+
[`${Provider.ANTHROPIC}:${Model.Anthropic.HAIKU_5_5}`]: {
|
|
1436
|
+
inputPricePerToken: 0.1 / PER_MILLION,
|
|
1437
|
+
outputPricePerToken: 0.5 / PER_MILLION
|
|
1438
|
+
},
|
|
1431
1439
|
[`${Provider.ANTHROPIC}:${Model.Anthropic.HAIKU_4_5}`]: {
|
|
1432
1440
|
inputPricePerToken: 1 / PER_MILLION,
|
|
1433
1441
|
outputPricePerToken: 5 / PER_MILLION
|