visual-ai-assertions 0.14.0 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # visual-ai-assertions
2
2
 
3
- AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, or Gemini and get structured, typed results.
3
+ AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, and Qwen via OpenRouter — and get structured, typed results.
4
4
 
5
5
  ## Installation
6
6
 
@@ -11,6 +11,7 @@ npm install visual-ai-assertions
11
11
  # Optional: install additional provider SDKs
12
12
  npm install @anthropic-ai/sdk # for Claude
13
13
  npm install @google/genai # for Gemini
14
+ # OpenRouter (Grok, Kimi, Qwen, ...) uses the OpenAI SDK — no extra install
14
15
 
15
16
  # Zod is a peer dependency
16
17
  npm install zod
@@ -441,11 +442,12 @@ The `VisualAIKnownError` union and `isVisualAIKnownError()` helper are useful wh
441
442
 
442
443
  ### API Keys
443
444
 
444
- | Provider | Environment Variable |
445
- | --------- | -------------------- |
446
- | Anthropic | `ANTHROPIC_API_KEY` |
447
- | OpenAI | `OPENAI_API_KEY` |
448
- | Google | `GOOGLE_API_KEY` |
445
+ | Provider | Environment Variable |
446
+ | ---------- | -------------------- |
447
+ | Anthropic | `ANTHROPIC_API_KEY` |
448
+ | OpenAI | `OPENAI_API_KEY` |
449
+ | Google | `GOOGLE_API_KEY` |
450
+ | OpenRouter | `OPENROUTER_API_KEY` |
449
451
 
450
452
  ### Optional Configuration
451
453
 
@@ -498,11 +500,12 @@ type SupportedMimeType = "image/jpeg" | "image/png" | "image/webp" | "image/gif"
498
500
 
499
501
  **Default models:**
500
502
 
501
- | Provider | Default Model |
502
- | --------- | ------------------------ |
503
- | Anthropic | `claude-sonnet-4-6` |
504
- | OpenAI | `gpt-5-mini` |
505
- | Google | `gemini-3-flash-preview` |
503
+ | Provider | Default Model |
504
+ | ---------- | ------------------------ |
505
+ | Anthropic | `claude-sonnet-4-6` |
506
+ | OpenAI | `gpt-5.4-mini` |
507
+ | Google | `gemini-3-flash-preview` |
508
+ | OpenRouter | `qwen/qwen3.6-flash` |
506
509
 
507
510
  ## Reasoning Effort
508
511
 
@@ -521,7 +524,8 @@ When omitted, each provider uses its default behavior. The `"xhigh"` level enabl
521
524
  | Anthropic (Fable 5/Opus 4.8/4.7/Sonnet 5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
522
525
  | Anthropic (other) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
523
526
  | OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
524
- | Google | `thinkingConfig.thinkingBudget` (1024 / 8192 / 24576) | `24576` (max budget) |
527
+ | Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
528
+ | OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
525
529
 
526
530
  ## Supported Models
527
531
 
@@ -544,8 +548,8 @@ All listed models support image/vision input. Pass any model ID to the `model` c
544
548
  | Model | Model ID | Input $/MTok | Output $/MTok | Notes |
545
549
  | ------------- | --------------- | ------------ | ------------- | --------------------------------- |
546
550
  | GPT-5.6 Sol | `gpt-5.6-sol` | $5 | $30 | Newest flagship, frontier tier |
547
- | GPT-5.6 Terra | `gpt-5.6-terra` | $2.50 | $15 | Newest balanced, everyday tier |
548
- | GPT-5.6 Luna | `gpt-5.6-luna` | $1 | $6 | Newest, fastest/cheapest tier |
551
+ | GPT-5.6 Terra | `gpt-5.6-terra` | $2 | $12 | Newest balanced, everyday tier |
552
+ | GPT-5.6 Luna | `gpt-5.6-luna` | $0.20 | $1.20 | Newest, fastest/cheapest tier |
549
553
  | GPT-5.5 | `gpt-5.5` | $5 | $30 | Previous flagship, 1M context |
550
554
  | GPT-5.4 Pro | `gpt-5.4-pro` | $30 | $180 | Most capable, extended context |
551
555
  | GPT-5.4 | `gpt-5.4` | $2.50 | $15 | Best vision quality |
@@ -558,11 +562,33 @@ All listed models support image/vision input. Pass any model ID to the `model` c
558
562
 
559
563
  | Model | Model ID | Input $/MTok | Output $/MTok | Notes |
560
564
  | --------------------- | ------------------------ | ------------ | ------------- | --------------------------------- |
565
+ | Gemini 3.8 Flash | `gemini-3.8-flash` | $0.75 | $3.75 | Newest GA flash; intro pricing¹ |
566
+ | Gemini 3.7 Flash | `gemini-3.7-flash` | $0.75 | $3.75 | Prior GA flash; intro pricing¹ |
567
+ | Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $7.50 | Prior GA flash; fewer out-tokens |
561
568
  | Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $9 | Strongest agentic & coding model |
569
+ | Gemini 3.5 Flash Lite | `gemini-3.5-flash-lite` | $0.30 | $2.50 | GA — fast, cheap, agentic tier |
562
570
  | Gemini 3.1 Pro | `gemini-3.1-pro-preview` | $2 | $12 | Preview — most advanced reasoning |
563
571
  | Gemini 3.1 Flash Lite | `gemini-3.1-flash-lite` | $0.25 | $1.50 | GA — lightweight and cheap |
564
572
  | Gemini 3 Flash | `gemini-3-flash-preview` | $0.50 | $3 | **Default** — fast and capable |
565
573
 
574
+ ¹ Gemini 3.8 Flash and 3.7 Flash introductory pricing runs through 2026-12-31; both revert to $1.50 / $7.50 per MTok on 2027-01-01.
575
+
576
+ ### OpenRouter
577
+
578
+ Any [OpenRouter](https://openrouter.ai/models) model slug (always `vendor/model`) is accepted — the vendor prefix is how the library recognizes an OpenRouter model. The models below are tested and have pricing built in. Note that OpenRouter may route a request to different upstream hosts with different quantizations; keep that in mind when comparing benchmark numbers.
579
+
580
+ | Model | Model ID | Input $/MTok | Output $/MTok | Notes |
581
+ | -------------- | --------------------------- | ------------ | ------------- | ------------------------------------- |
582
+ | Grok 4.6 | `x-ai/grok-4.6` | $2 | $6 | Newest xAI flagship, 500K context |
583
+ | Grok 4.5 | `x-ai/grok-4.5` | $2 | $6 | Prior xAI flagship, 500K context |
584
+ | Kimi K3 | `moonshotai/kimi-k3` | $3 | $15 | Moonshot flagship, 1M context |
585
+ | Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | $0.82 | $3.75 | Agentic/coding tier with vision |
586
+ | Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input |
587
+ | Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading |
588
+ | Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 | **Default** — cheap flash vision tier |
589
+
590
+ `qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
591
+
566
592
  ## License
567
593
 
568
594
  MIT