visual-ai-assertions 0.26.0 → 0.27.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,6 +2,8 @@
2
2
 
3
3
  AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, Qwen, and GLM via OpenRouter — and get structured, typed results.
4
4
 
5
+ Every supported model is benchmarked against hand-labelled screenshots — **[see how they compare](https://nullp2ike.github.io/visual-reasoning/)** ([methodology](#benchmarks)).
6
+
5
7
  ## Installation
6
8
 
7
9
  ```bash
@@ -589,14 +591,14 @@ Set `VISUAL_AI_REASONING_EFFORT` to apply one without touching code; an explicit
589
591
 
590
592
  When omitted, each provider uses its default behavior. The `"xhigh"` level enables maximum reasoning depth. `"minimal"` is **not portable** — several OpenAI models reject it with HTTP 400 (including the `gpt-6.1-sol` default), and Google/OpenRouter clamp it to `low` rather than send it; use `"low"` as the floor unless you know the model accepts it.
591
593
 
592
- | Provider | Native Parameter | `"xhigh"` maps to |
593
- | --------------------------------------------------------- | ----------------------------------------------------- | -------------------- |
594
- | Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
595
- | Anthropic (Opus 4.6, Sonnet 4.6) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
596
- | Anthropic (Haiku 4.5) | budget-based extended thinking (token budget) | 16384-token budget |
597
- | OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
598
- | Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
599
- | OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
594
+ | Provider | Native Parameter | `"xhigh"` maps to |
595
+ | -------------------------------------------------------------------- | ----------------------------------------------------- | -------------------- |
596
+ | Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5, Haiku 5.5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
597
+ | Anthropic (Opus 4.6, Sonnet 4.6) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
598
+ | Anthropic (Haiku 4.5) | budget-based extended thinking (token budget) | 16384-token budget |
599
+ | OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
600
+ | Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
601
+ | OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
600
602
 
601
603
  ## Supported Models
602
604
 
@@ -604,19 +606,22 @@ All listed models support image/vision input. Pass any model ID to the `model` c
604
606
 
605
607
  ### Anthropic
606
608
 
607
- | Model | Model ID | Input $/MTok | Output $/MTok | Notes |
608
- | ----------------- | ------------------- | ------------ | ------------- | --------------------------------------------- |
609
- | Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work |
610
- | Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price |
611
- | Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5 |
612
- | Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh` |
613
- | Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh` |
614
- | Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier |
615
- | Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output |
616
- | Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh` |
617
- | Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work |
618
- | Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation |
619
- | Claude Haiku 4.5 | `claude-haiku-4-5` | $1 | $5 | Fastest, budget-friendly |
609
+ | Model | Model ID | Input $/MTok | Output $/MTok | Notes |
610
+ | ----------------- | ------------------- | ------------ | ------------- | ------------------------------------------------- |
611
+ | Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work |
612
+ | Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price |
613
+ | Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5 |
614
+ | Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh` |
615
+ | Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh` |
616
+ | Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier |
617
+ | Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output |
618
+ | Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh` |
619
+ | Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work |
620
+ | Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation |
621
+ | Claude Haiku 5.5 | `claude-haiku-5-5` | $0.10 | $0.50 | Newest Haiku; cheapest Claude; supports `xhigh` † |
622
+ | Claude Haiku 4.5 | `claude-haiku-4-5` | $1 | $5 | Fastest, budget-friendly |
623
+
624
+ † Haiku 5.5 prices apply to prompts up to 100K tokens; longer prompts bill at $0.50 / $2.50, far beyond screenshot-sized calls. Unlike Haiku 4.5 it uses adaptive thinking like the other current Claude models, and it thinks even when `reasoningEffort` is not set.
620
625
 
621
626
  ### OpenAI
622
627
 
@@ -689,6 +694,51 @@ Meta also publishes `meta/muse-spark-1.3-contributor`, the same model at $0.10 /
689
694
 
690
695
  `qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
691
696
 
697
+ ## Benchmarks
698
+
699
+ Which model should you actually pick? This repo benchmarks every supported model
700
+ against hand-labelled screenshots and publishes the results:
701
+
702
+ **📊 [Live results — nullp2ike.github.io/visual-reasoning](https://nullp2ike.github.io/visual-reasoning/)**
703
+
704
+ The site carries the defect-discovery leaderboard, a screenshot × model matrix you
705
+ can drill into for any individual answer, and the same runs graded independently by
706
+ five different judges so you can see where the grading itself is contested.
707
+
708
+ The benchmark asks each model one open question — "What looks visually broken on
709
+ this page?" — with no hints, and an LLM judge matches what it reports against the
710
+ defects seeded into each screenshot. The headline is recall of those defects,
711
+ weighed against the extras a model invents per run: a missed defect is a bug that
712
+ ships, and an invented one is noise someone has to triage.
713
+
714
+ ### The `golden` dataset
715
+
716
+ 18 screenshots — 17 with exactly one deliberately seeded defect, plus a clean
717
+ control where anything reported counts as a false positive. The prompt names a
718
+ short list of out-of-scope non-defects (edge-clipped carousel items, content cut
719
+ off by the viewport bottom) so models are not penalised for reporting framing as
720
+ breakage.
721
+
722
+ Current coverage, at `medium` effort with 5 repeats per cell:
723
+
724
+ - 38 model/effort/fidelity series over 3,420 graded runs
725
+
726
+ A few results worth knowing before you choose a default:
727
+
728
+ | Model | Discovery recall | Extras / run | Cost / run |
729
+ | ----------------------------------------- | ---------------- | ------------ | ---------- |
730
+ | `claude-opus-5-5` | 100% | 0.53 | $0.023 |
731
+ | `claude-sonnet-5-5` _(Anthropic default)_ | 94% | 2.81 | $0.011 |
732
+ | `gpt-6.1-sol` _(OpenAI default)_ | 93% | 0.02 | $0.0055 |
733
+ | `gemini-3.8-flash` _(Google default)_ | 87% | 0.34 | $0.0050 |
734
+
735
+ Recall is not the whole picture. `claude-sonnet-5-5` matches the Fable tier on
736
+ recall but reports 2.81 extras per run against `gpt-6.1-sol`'s 0.02 — for
737
+ open-ended discovery that is a lot of noise to triage.
738
+
739
+ Full methodology, how to run a sweep, and how to add your own dataset:
740
+ [`bench/README.md`](bench/README.md).
741
+
692
742
  ## License
693
743
 
694
744
  MIT
package/dist/index.cjs CHANGED
@@ -497,6 +497,7 @@ var Model = {
497
497
  SONNET_5_5: "claude-sonnet-5-5",
498
498
  SONNET_5: "claude-sonnet-5",
499
499
  SONNET_4_6: "claude-sonnet-4-6",
500
+ HAIKU_5_5: "claude-haiku-5-5",
500
501
  HAIKU_4_5: "claude-haiku-4-5"
501
502
  },
502
503
  OpenAI: {
@@ -716,7 +717,8 @@ var XHIGH_CAPABLE_MODELS = /* @__PURE__ */ new Set([
716
717
  Model.Anthropic.OPUS_4_8,
717
718
  Model.Anthropic.OPUS_4_7,
718
719
  Model.Anthropic.SONNET_5_5,
719
- Model.Anthropic.SONNET_5
720
+ Model.Anthropic.SONNET_5,
721
+ Model.Anthropic.HAIKU_5_5
720
722
  ]);
721
723
  function mapEffort(level, model) {
722
724
  if (level === "minimal") return "low";
@@ -1428,6 +1430,12 @@ var PRICING_TABLE = {
1428
1430
  inputPricePerToken: 3 / PER_MILLION,
1429
1431
  outputPricePerToken: 15 / PER_MILLION
1430
1432
  },
1433
+ // Prompts above 100K tokens bill at $0.50/$2.50, far beyond screenshot-sized
1434
+ // calls. Cached input is $0.01/MTok and cache writes $0.125/MTok (not modelled).
1435
+ [`${Provider.ANTHROPIC}:${Model.Anthropic.HAIKU_5_5}`]: {
1436
+ inputPricePerToken: 0.1 / PER_MILLION,
1437
+ outputPricePerToken: 0.5 / PER_MILLION
1438
+ },
1431
1439
  [`${Provider.ANTHROPIC}:${Model.Anthropic.HAIKU_4_5}`]: {
1432
1440
  inputPricePerToken: 1 / PER_MILLION,
1433
1441
  outputPricePerToken: 5 / PER_MILLION