visual-ai-assertions 0.23.0 → 0.26.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +140 -83
- package/dist/index.cjs +600 -190
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +151 -30
- package/dist/index.d.ts +151 -30
- package/dist/index.js +600 -190
- package/dist/index.js.map +1 -1
- package/package.json +8 -9
package/README.md
CHANGED
|
@@ -99,7 +99,7 @@ const ai = visualAI();
|
|
|
99
99
|
|
|
100
100
|
// Explicit configuration
|
|
101
101
|
const ai = visualAI({
|
|
102
|
-
model: "claude-sonnet-
|
|
102
|
+
model: "claude-sonnet-5-5", // optional, sensible defaults per provider
|
|
103
103
|
apiKey: "sk-...", // optional, defaults to provider env var
|
|
104
104
|
debug: true, // optional, logs prompts/responses to stderr
|
|
105
105
|
maxTokens: 4096, // optional, default 4096
|
|
@@ -110,7 +110,7 @@ const ai = visualAI({
|
|
|
110
110
|
|
|
111
111
|
// Use constants for IDE autocomplete
|
|
112
112
|
const ai = visualAI({
|
|
113
|
-
model: Model.Anthropic.
|
|
113
|
+
model: Model.Anthropic.SONNET_5_5,
|
|
114
114
|
});
|
|
115
115
|
```
|
|
116
116
|
|
|
@@ -198,8 +198,8 @@ import { writeFileSync } from "node:fs";
|
|
|
198
198
|
// Basic comparison
|
|
199
199
|
const result = await ai.compare(before, after);
|
|
200
200
|
|
|
201
|
-
// gemini-3-flash
|
|
202
|
-
// Pass { diffImage: false } to opt out.
|
|
201
|
+
// Every Gemini flash model — the gemini-3.8-flash default included — auto-includes
|
|
202
|
+
// an annotated diff. Pass { diffImage: false } to opt out.
|
|
203
203
|
|
|
204
204
|
// With custom prompt and instructions
|
|
205
205
|
const result = await ai.compare(before, after, {
|
|
@@ -207,8 +207,9 @@ const result = await ai.compare(before, after, {
|
|
|
207
207
|
instructions: ["Ignore date/time differences"],
|
|
208
208
|
});
|
|
209
209
|
|
|
210
|
-
//
|
|
211
|
-
//
|
|
210
|
+
// Requesting the diff image explicitly (it is already on by default for the flash tier).
|
|
211
|
+
// Supported on gemini-3-flash-preview, 3.5, 3.6, 3.7 and 3.8-flash (DIFF_ALLOWED_MODELS);
|
|
212
|
+
// Flash-Lite and Pro tiers produce none even when asked.
|
|
212
213
|
const result = await ai.compare(before, after, {
|
|
213
214
|
diffImage: true,
|
|
214
215
|
});
|
|
@@ -224,7 +225,7 @@ if (result.diffImage) {
|
|
|
224
225
|
pass: boolean; // true if no critical/major changes
|
|
225
226
|
reasoning: string; // overall summary
|
|
226
227
|
changes: ChangeEntry[]; // list of visual differences
|
|
227
|
-
diffImage?: { // present when diffing is enabled explicitly or by Gemini
|
|
228
|
+
diffImage?: { // present when diffing is enabled explicitly or by the Gemini flash default
|
|
228
229
|
data: Buffer; // PNG image data
|
|
229
230
|
width: number;
|
|
230
231
|
height: number;
|
|
@@ -352,16 +353,22 @@ Oversized images are automatically resized to provider limits.
|
|
|
352
353
|
|
|
353
354
|
`ai.check()` and `ai.ask()` also accept short video recordings (`.mp4`, `.webm`, `.mov`, `.mkv`) — useful for asserting on transient UI like toast messages. Accepted shapes are file path, `data:video/...;base64,...` URL, raw base64 string, `Buffer`, and `Uint8Array`. HTTP/HTTPS URLs are not supported for video inputs — fetch the bytes yourself first.
|
|
354
355
|
|
|
356
|
+
How the video reaches the model depends on the provider. **Google models receive the video itself** (native delivery): Gemini samples and tokenises it server-side and also hears the audio track. **Every other provider gets frames sampled with ffmpeg** and sent as an ordered image timeline. Both paths enforce the same duration cap before any provider call, and both give you per-statement timestamps.
|
|
357
|
+
|
|
355
358
|
```typescript
|
|
356
359
|
// Playwright recording on disk
|
|
357
360
|
const result = await ai.check("./trace/video/recording.webm", [
|
|
358
361
|
'A success toast with text "Saved" briefly appears',
|
|
359
362
|
]);
|
|
363
|
+
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
364
|
+
|
|
365
|
+
// Google model → the video was sent natively
|
|
366
|
+
console.log(result.video);
|
|
367
|
+
// { durationSeconds: 4.0, fps: 1, mimeType: "video/webm", delivery: "inline" }
|
|
360
368
|
|
|
361
|
-
//
|
|
369
|
+
// Any other provider → frames were sampled
|
|
362
370
|
console.log(result.frames);
|
|
363
|
-
// { count:
|
|
364
|
-
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
371
|
+
// { count: 3, timestampsSeconds: [0.5, 2.5, 3.5], durationSeconds: 4.0, droppedUnchanged: 1 }
|
|
365
372
|
|
|
366
373
|
// Override sampling — defaults are 1 fps, max 10 frames, max 10 s of video
|
|
367
374
|
await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
@@ -369,9 +376,37 @@ await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
|
369
376
|
});
|
|
370
377
|
```
|
|
371
378
|
|
|
372
|
-
`maxFrames` is hard-capped at 60 to keep memory bounded
|
|
379
|
+
`fps` applies to both paths: it is the ffmpeg sampling rate, and for native delivery it is passed to the provider as its sampling rate. `maxFrames` and `dedupe` apply to frame sampling only; `maxFrames` is hard-capped at 60 to keep memory bounded, and frames are downscaled so the longer edge fits within 1568 px before being sent.
|
|
380
|
+
|
|
381
|
+
**Choosing the delivery.** `video.mode` is `"auto"` by default: native where the provider supports it, frames elsewhere. Set `"frames"` to sample frames on a Google model too, or `"native"` to insist on native delivery — that throws `VisualAIConfigError` on providers without it.
|
|
382
|
+
|
|
383
|
+
```typescript
|
|
384
|
+
// Sample frames on Gemini instead of sending the video
|
|
385
|
+
await ai.check("./clip.mp4", ["Loader disappears"], { video: { mode: "frames" } });
|
|
386
|
+
|
|
387
|
+
// Fail loudly if the configured model cannot take video natively
|
|
388
|
+
await ai.ask("./clip.mp4", "What happens?", { video: { mode: "native" } });
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
Native delivery is worth preferring where available. Measured on a 15 s recording with a seven-statement `check()` on Gemini 3.8 Flash, native video at 1 fps cost about 60% of the frames path with dedupe on and answered about 1.5 s faster, with identical verdicts and correct timestamps on every run; Gemini tokenises a video frame at roughly 60 tokens against about 1,080 for the same frame sent as an image. Videos over 19 MB are uploaded through the Gemini Files API automatically (`delivery: "file"`) and deleted again afterwards. For `ask()`, native results carry `timestampReferences` (seconds) in place of the frame-indexed `frameReferences`.
|
|
373
392
|
|
|
374
|
-
|
|
393
|
+
**Unchanged frames are dropped.** A recording of a mostly static screen would otherwise pay input tokens for every near-identical sample, so by default each sampled frame is compared against the most recently kept frame and dropped when fewer than 0.1% of its pixels changed — roughly a 37x37 px region on a 1568x880 frame. A toast, a loader, or a dialog comfortably clears that; compression noise and a blinking text caret do not. The first frame is always kept, `durationSeconds` still covers the whole clip, and the prompt tells the model how many frames were dropped and that the screen stayed unchanged between the listed timestamps. `result.frames.droppedUnchanged` reports how many were dropped.
|
|
394
|
+
|
|
395
|
+
```typescript
|
|
396
|
+
// Keep every sampled frame
|
|
397
|
+
await ai.check("./clip.webm", ["A spinner is visible throughout"], {
|
|
398
|
+
video: { dedupe: false },
|
|
399
|
+
});
|
|
400
|
+
|
|
401
|
+
// Keep smaller changes (fraction of pixels that must change, in (0, 1])
|
|
402
|
+
await ai.check("./clip.webm", ["The 24 px status dot turns green"], {
|
|
403
|
+
video: { dedupe: { threshold: 0.0002 } },
|
|
404
|
+
});
|
|
405
|
+
```
|
|
406
|
+
|
|
407
|
+
A change smaller than the threshold — a lone 24x24 px icon on a full-size frame is about 0.04% — is treated as no change, so lower `threshold` or pass `dedupe: false` when the assertion is about something that small.
|
|
408
|
+
|
|
409
|
+
How it works: on the frames path the library samples frames with ffmpeg, drops the ones that did not visibly change, and sends the rest to the provider as an ordered timeline; on the native path it probes the duration, then hands the provider the video and a matching prompt. Either way a statement passes when it is true at any point, unless its wording specifies otherwise (e.g. "throughout"). Template helpers (`accessibility`, `layout`, `pageLoad`, `content`, `elementsVisible`, `elementsHidden`) are image-only — pass video to `check()` or `ask()` instead.
|
|
375
410
|
|
|
376
411
|
**ffmpeg setup.** Video support works out of the box — `fluent-ffmpeg`, `@ffmpeg-installer/ffmpeg`, and `@ffprobe-installer/ffprobe` ship as regular dependencies and bundle platform-specific ffmpeg/ffprobe binaries. If you ran `npm install` you already have everything you need. On platforms where the prebuilt binary is unavailable (or if you've pruned dependencies), `check()` and `ask()` throw `VisualAIVideoError` (import from `visual-ai-assertions` to `instanceof`-narrow it) when called with video input.
|
|
377
412
|
|
|
@@ -384,7 +419,7 @@ const frames = [await page.screenshot(), await page.screenshot()];
|
|
|
384
419
|
|
|
385
420
|
const result = await ai.check({ frames }, ['A success toast with text "Saved" appears']);
|
|
386
421
|
console.log(result.frames);
|
|
387
|
-
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1 }
|
|
422
|
+
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1, droppedUnchanged: 0 }
|
|
388
423
|
|
|
389
424
|
// Control timestamps: give an explicit fps, or per-frame timestampSeconds
|
|
390
425
|
await ai.ask(
|
|
@@ -398,7 +433,7 @@ await ai.ask(
|
|
|
398
433
|
);
|
|
399
434
|
```
|
|
400
435
|
|
|
401
|
-
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path.
|
|
436
|
+
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path — pre-sampled frames are always sent as frames, even to Google models. Frames that did not visibly change from the preceding kept frame are dropped exactly as for video input; control it with `dedupe` on the `FramesInput` itself, e.g. `{ frames, dedupe: false }` or `{ frames, dedupe: { threshold: 0.0002 } }`.
|
|
402
437
|
|
|
403
438
|
### Formatting & Assertion Helpers
|
|
404
439
|
|
|
@@ -479,15 +514,16 @@ The `VisualAIKnownError` union and `isVisualAIKnownError()` helper are useful wh
|
|
|
479
514
|
|
|
480
515
|
### Optional Configuration
|
|
481
516
|
|
|
482
|
-
| Variable | Description
|
|
483
|
-
| ---------------------------- |
|
|
484
|
-
| `VISUAL_AI_MODEL` |
|
|
485
|
-
| `
|
|
486
|
-
| `
|
|
487
|
-
| `
|
|
488
|
-
| `
|
|
489
|
-
| `
|
|
490
|
-
| `
|
|
517
|
+
| Variable | Description |
|
|
518
|
+
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
519
|
+
| `VISUAL_AI_MODEL` | Model used when `model` is not set in config, overriding the provider's default. It also **selects the provider**, since the provider is inferred from the model name — so this is how you pick a provider when several API keys are set. |
|
|
520
|
+
| `VISUAL_AI_REASONING_EFFORT` | Reasoning effort when `reasoningEffort` is not set in config. One of `minimal`, `low`, `medium`, `high`, `xhigh` (case-insensitive). An unrecognised value throws `VisualAIConfigError` rather than being ignored. |
|
|
521
|
+
| `VISUAL_AI_DEBUG` | Enable error diagnostic logging to stderr. Does **not** enable prompt/response logging. Use `"true"` or `"1"`. |
|
|
522
|
+
| `VISUAL_AI_DEBUG_PROMPT` | Enable prompt-only debug logging to stderr. Use `"true"` or `"1"`. |
|
|
523
|
+
| `VISUAL_AI_DEBUG_RESPONSE` | Enable response-only debug logging to stderr. Use `"true"` or `"1"`. |
|
|
524
|
+
| `VISUAL_AI_DEBUG_FRAMES` | Persist sampled video frames to disk for offline inspection. Use `"true"` or `"1"`. Frames are written to `./visual-ai-debug-frames/<timestamp>-<id>/` (override path with the next variable). Has no effect on image-only inputs. |
|
|
525
|
+
| `VISUAL_AI_DEBUG_FRAMES_DIR` | Override the base directory for `VISUAL_AI_DEBUG_FRAMES`. Each call still gets its own timestamped subdirectory inside it. |
|
|
526
|
+
| `VISUAL_AI_TRACK_USAGE` | Enable usage tracking (token counts and cost) to stderr. Use `"true"` or `"1"`. |
|
|
491
527
|
|
|
492
528
|
## Configuration
|
|
493
529
|
|
|
@@ -511,7 +547,10 @@ import type {
|
|
|
511
547
|
CheckResult,
|
|
512
548
|
CompareResult,
|
|
513
549
|
Frame,
|
|
550
|
+
FrameDedupeOptions,
|
|
514
551
|
MediaInput,
|
|
552
|
+
NativeVideoMetadata,
|
|
553
|
+
VideoDeliveryMode,
|
|
515
554
|
SupportedMimeType,
|
|
516
555
|
SupportedVideoMimeType,
|
|
517
556
|
VideoFramesMetadata,
|
|
@@ -529,12 +568,12 @@ type SupportedMimeType = "image/jpeg" | "image/png" | "image/webp" | "image/gif"
|
|
|
529
568
|
|
|
530
569
|
**Default models:**
|
|
531
570
|
|
|
532
|
-
| Provider | Default Model
|
|
533
|
-
| ---------- |
|
|
534
|
-
| Anthropic | `claude-sonnet-
|
|
535
|
-
| OpenAI | `gpt-
|
|
536
|
-
| Google | `gemini-3-flash
|
|
537
|
-
| OpenRouter | `
|
|
571
|
+
| Provider | Default Model |
|
|
572
|
+
| ---------- | --------------------- |
|
|
573
|
+
| Anthropic | `claude-sonnet-5-5` |
|
|
574
|
+
| OpenAI | `gpt-6.1-sol` |
|
|
575
|
+
| Google | `gemini-3.8-flash` |
|
|
576
|
+
| OpenRouter | `meta/muse-spark-1.3` |
|
|
538
577
|
|
|
539
578
|
## Reasoning Effort
|
|
540
579
|
|
|
@@ -542,19 +581,22 @@ Control how deeply the model reasons before responding. Higher effort produces m
|
|
|
542
581
|
|
|
543
582
|
```typescript
|
|
544
583
|
const ai = visualAI({
|
|
545
|
-
reasoningEffort: "high", // "low" | "medium" | "high" | "xhigh"
|
|
584
|
+
reasoningEffort: "high", // "minimal" | "low" | "medium" | "high" | "xhigh"
|
|
546
585
|
});
|
|
547
586
|
```
|
|
548
587
|
|
|
549
|
-
|
|
588
|
+
Set `VISUAL_AI_REASONING_EFFORT` to apply one without touching code; an explicit `reasoningEffort` wins over it.
|
|
550
589
|
|
|
551
|
-
|
|
552
|
-
|
|
553
|
-
|
|
|
554
|
-
|
|
|
555
|
-
|
|
|
556
|
-
|
|
|
557
|
-
|
|
|
590
|
+
When omitted, each provider uses its default behavior. The `"xhigh"` level enables maximum reasoning depth. `"minimal"` is **not portable** — several OpenAI models reject it with HTTP 400 (including the `gpt-6.1-sol` default), and Google/OpenRouter clamp it to `low` rather than send it; use `"low"` as the floor unless you know the model accepts it.
|
|
591
|
+
|
|
592
|
+
| Provider | Native Parameter | `"xhigh"` maps to |
|
|
593
|
+
| --------------------------------------------------------- | ----------------------------------------------------- | -------------------- |
|
|
594
|
+
| Anthropic (Fable 5.1/5, Opus 5.5/5/4.8/4.7, Sonnet 5.5/5) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "xhigh"` |
|
|
595
|
+
| Anthropic (Opus 4.6, Sonnet 4.6) | `thinking.type: "adaptive"` + `output_config.effort` | `effort: "max"` |
|
|
596
|
+
| Anthropic (Haiku 4.5) | budget-based extended thinking (token budget) | 16384-token budget |
|
|
597
|
+
| OpenAI | `reasoning.effort` (Responses API) | `effort: "xhigh"` |
|
|
598
|
+
| Google | `thinkingConfig.thinkingLevel` (1:1: low/medium/high) | `"high"` (max level) |
|
|
599
|
+
| OpenRouter | `reasoning.effort` (normalized low/medium/high) | `effort: "high"` |
|
|
558
600
|
|
|
559
601
|
## Supported Models
|
|
560
602
|
|
|
@@ -562,49 +604,57 @@ All listed models support image/vision input. Pass any model ID to the `model` c
|
|
|
562
604
|
|
|
563
605
|
### Anthropic
|
|
564
606
|
|
|
565
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
566
|
-
| ----------------- | ------------------- | ------------ | ------------- |
|
|
567
|
-
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work
|
|
568
|
-
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price
|
|
569
|
-
| Claude Opus
|
|
570
|
-
| Claude Opus
|
|
571
|
-
| Claude Opus 4.
|
|
572
|
-
| Claude
|
|
573
|
-
| Claude
|
|
574
|
-
| Claude
|
|
607
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
608
|
+
| ----------------- | ------------------- | ------------ | ------------- | --------------------------------------------- |
|
|
609
|
+
| Claude Fable 5.1 | `claude-fable-5-1` | $10 | $50 | Most capable; long-horizon agentic work |
|
|
610
|
+
| Claude Fable 5 | `claude-fable-5` | $10 | $50 | Predecessor to Fable 5.1, same price |
|
|
611
|
+
| Claude Opus 5.5 | `claude-opus-5-5` | $4 | $20 | Newest Opus; >30% faster output than Opus 5 |
|
|
612
|
+
| Claude Opus 5 | `claude-opus-5` | $5 | $25 | Previous Opus flagship; supports `xhigh` |
|
|
613
|
+
| Claude Opus 4.8 | `claude-opus-4-8` | $5 | $25 | Prior Opus tier; supports `xhigh` |
|
|
614
|
+
| Claude Opus 4.7 | `claude-opus-4-7` | $5 | $25 | Previous Opus; supports `xhigh` effort tier |
|
|
615
|
+
| Claude Opus 4.6 | `claude-opus-4-6` | $5 | $25 | Previous flagship, 128K max output |
|
|
616
|
+
| Claude Sonnet 5.5 | `claude-sonnet-5-5` | $2 | $10 | **Default** — newest Sonnet; supports `xhigh` |
|
|
617
|
+
| Claude Sonnet 5 | `claude-sonnet-5` | $3 | $15 | Near-Opus quality on coding/agentic work |
|
|
618
|
+
| Claude Sonnet 4.6 | `claude-sonnet-4-6` | $3 | $15 | Prior default; best value in its generation |
|
|
619
|
+
| Claude Haiku 4.5 | `claude-haiku-4-5` | $1 | $5 | Fastest, budget-friendly |
|
|
575
620
|
|
|
576
621
|
### OpenAI
|
|
577
622
|
|
|
578
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
579
|
-
| ------------- | --------------- | ------------ | ------------- |
|
|
580
|
-
| GPT-6 Astra | `gpt-6-astra` | $10 | $50 | Most capable; restricted access¹
|
|
581
|
-
| GPT-
|
|
582
|
-
| GPT-
|
|
583
|
-
| GPT-
|
|
584
|
-
| GPT-5.
|
|
585
|
-
| GPT-5.
|
|
586
|
-
| GPT-5.
|
|
587
|
-
| GPT-5.
|
|
588
|
-
| GPT-5.4
|
|
589
|
-
| GPT-5.4
|
|
590
|
-
| GPT-5
|
|
623
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
624
|
+
| ------------- | --------------- | ------------ | ------------- | ----------------------------------- |
|
|
625
|
+
| GPT-6 Astra | `gpt-6-astra` | $10 | $50 | Most capable; restricted access¹ |
|
|
626
|
+
| GPT-6.1 Sol | `gpt-6.1-sol` | $2 | $10 | **Default** — near-Astra quality² |
|
|
627
|
+
| GPT-6 Sol | `gpt-6-sol` | $2 | $10 | GPT-6 generation, frontier tier |
|
|
628
|
+
| GPT-6 Luna | `gpt-6-luna` | $0.10 | $0.50 | GPT-6 generation, fastest/cheapest |
|
|
629
|
+
| GPT-5.6 Sol | `gpt-5.6-sol` | $5 | $30 | Previous flagship, frontier tier |
|
|
630
|
+
| GPT-5.6 Terra | `gpt-5.6-terra` | $2 | $12 | Newest balanced, everyday tier |
|
|
631
|
+
| GPT-5.6 Luna | `gpt-5.6-luna` | $0.20 | $1.20 | Prior default — fast and cheap |
|
|
632
|
+
| GPT-5.5 | `gpt-5.5` | $5 | $30 | Previous flagship, 1M context |
|
|
633
|
+
| GPT-5.4 Pro | `gpt-5.4-pro` | $30 | $180 | Most capable, extended context |
|
|
634
|
+
| GPT-5.4 | `gpt-5.4` | $2.50 | $15 | Best vision quality |
|
|
635
|
+
| GPT-5.2 | `gpt-5.2` | $1.75 | $14 | Balanced quality and cost |
|
|
636
|
+
| GPT-5.4 mini | `gpt-5.4-mini` | $0.75 | $4.50 | Prior default — fast and affordable |
|
|
637
|
+
| GPT-5.4 nano | `gpt-5.4-nano` | $0.20 | $1.25 | Cheapest older-generation option |
|
|
638
|
+
| GPT-5 mini | `gpt-5-mini` | $0.25 | $2 | Fast and cheap |
|
|
591
639
|
|
|
592
640
|
¹ GPT-6 Astra is rolling out through OpenAI's Trusted Access Program, so many API keys cannot reach it yet — expect a `VisualAIProviderError` naming the model until your account is enabled.
|
|
593
641
|
|
|
594
|
-
Astra
|
|
642
|
+
Astra uses the same output budget as other OpenAI models: the 4096 default, raised to 16384 automatically at `high`/`xhigh`. It used to get 32768 at every effort because plain `ask()` calls exhausted the default, which looked like heavy reasoning. That was actually the image `ask()` schema bug fixed in this release: the model spent under 50 reasoning tokens, then printed whitespace until the budget ran out. With the fix, `ask()` and `check()` both completed every call at 4096 in live testing, and no call on the `golden` bench produced more than 548 output tokens. It also accepts a fifth reasoning level, `max`, above `xhigh`; this library's `reasoningEffort` stops at `xhigh`, which is passed through unchanged, so `max` is not currently reachable.
|
|
643
|
+
|
|
644
|
+
² GPT-6.1 Sol accepts `low`, `medium`, `high`, `xhigh` and `max`; `minimal` and `none` return HTTP 400, so use `low` as the floor.
|
|
595
645
|
|
|
596
646
|
### Google
|
|
597
647
|
|
|
598
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
599
|
-
| --------------------- | ------------------------ | ------------ | ------------- |
|
|
600
|
-
| Gemini 3.8 Flash | `gemini-3.8-flash` | $0.75 | $3.75 |
|
|
601
|
-
| Gemini 3.7 Flash | `gemini-3.7-flash` | $0.75 | $3.75 | Prior GA flash; intro pricing¹
|
|
602
|
-
| Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $7.50 | Prior GA flash; fewer out-tokens
|
|
603
|
-
| Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $9 | Strongest agentic & coding model
|
|
604
|
-
| Gemini 3.5 Flash Lite | `gemini-3.5-flash-lite` | $0.30 | $2.50 | GA — fast, cheap, agentic tier
|
|
605
|
-
| Gemini 3.1 Pro | `gemini-3.1-pro-preview` | $2 | $12 | Preview — most advanced reasoning
|
|
606
|
-
| Gemini 3.1 Flash Lite | `gemini-3.1-flash-lite` | $0.25 | $1.50 | GA — lightweight and cheap
|
|
607
|
-
| Gemini 3 Flash | `gemini-3-flash-preview` | $0.50 | $3 |
|
|
648
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
649
|
+
| --------------------- | ------------------------ | ------------ | ------------- | ---------------------------------- |
|
|
650
|
+
| Gemini 3.8 Flash | `gemini-3.8-flash` | $0.75 | $3.75 | **Default** — newest GA flash¹ |
|
|
651
|
+
| Gemini 3.7 Flash | `gemini-3.7-flash` | $0.75 | $3.75 | Prior GA flash; intro pricing¹ |
|
|
652
|
+
| Gemini 3.6 Flash | `gemini-3.6-flash` | $1.50 | $7.50 | Prior GA flash; fewer out-tokens |
|
|
653
|
+
| Gemini 3.5 Flash | `gemini-3.5-flash` | $1.50 | $9 | Strongest agentic & coding model |
|
|
654
|
+
| Gemini 3.5 Flash Lite | `gemini-3.5-flash-lite` | $0.30 | $2.50 | GA — fast, cheap, agentic tier |
|
|
655
|
+
| Gemini 3.1 Pro | `gemini-3.1-pro-preview` | $2 | $12 | Preview — most advanced reasoning |
|
|
656
|
+
| Gemini 3.1 Flash Lite | `gemini-3.1-flash-lite` | $0.25 | $1.50 | GA — lightweight and cheap |
|
|
657
|
+
| Gemini 3 Flash | `gemini-3-flash-preview` | $0.50 | $3 | Prior default; cheapest flash tier |
|
|
608
658
|
|
|
609
659
|
¹ Gemini 3.8 Flash and 3.7 Flash introductory pricing runs through 2026-12-31; both revert to $1.50 / $7.50 per MTok on 2027-01-01.
|
|
610
660
|
|
|
@@ -612,22 +662,29 @@ Astra reasons heavily enough to spend the entire 4096-token default output budge
|
|
|
612
662
|
|
|
613
663
|
Any [OpenRouter](https://openrouter.ai/models) model slug (always `vendor/model`) is accepted — the vendor prefix is how the library recognizes an OpenRouter model. The models below are tested and have pricing built in. Note that OpenRouter may route a request to different upstream hosts with different quantizations; keep that in mind when comparing benchmark numbers.
|
|
614
664
|
|
|
615
|
-
| Model | Model ID | Input $/MTok | Output $/MTok | Notes
|
|
616
|
-
| -------------- | --------------------------- | ------------ | ------------- |
|
|
617
|
-
| Muse Spark 1.3 | `meta/muse-spark-1.3` | $1.25 | $4.25 | Meta flagship
|
|
618
|
-
| Grok 4.6 | `x-ai/grok-4.6` | $2 | $6 | Newest xAI flagship, 500K context
|
|
619
|
-
| Grok 4.5 | `x-ai/grok-4.5` | $2 | $6 | Prior xAI flagship, 500K context
|
|
620
|
-
| Kimi K3 | `moonshotai/kimi-k3` | $3 | $15 | Moonshot flagship, 1M context
|
|
621
|
-
| Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | $0.82 | $3.75 | Agentic/coding tier with vision
|
|
622
|
-
| Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input
|
|
623
|
-
| Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading
|
|
624
|
-
| Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 |
|
|
625
|
-
| GLM 5.3 Flash | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | Z.ai flash tier, 1.3M context²
|
|
665
|
+
| Model | Model ID | Input $/MTok | Output $/MTok | Notes |
|
|
666
|
+
| -------------- | --------------------------- | ------------ | ------------- | ----------------------------------- |
|
|
667
|
+
| Muse Spark 1.3 | `meta/muse-spark-1.3` | $1.25 | $4.25 | **Default** — Meta flagship; gated¹ |
|
|
668
|
+
| Grok 4.6 | `x-ai/grok-4.6` | $2 | $6 | Newest xAI flagship, 500K context |
|
|
669
|
+
| Grok 4.5 | `x-ai/grok-4.5` | $2 | $6 | Prior xAI flagship, 500K context |
|
|
670
|
+
| Kimi K3 | `moonshotai/kimi-k3` | $3 | $15 | Moonshot flagship, 1M context |
|
|
671
|
+
| Kimi K2.7 Code | `moonshotai/kimi-k2.7-code` | $0.82 | $3.75 | Agentic/coding tier with vision⁵ |
|
|
672
|
+
| Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input⁴ |
|
|
673
|
+
| Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading⁴ |
|
|
674
|
+
| Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 | Prior default — cheap flash vision |
|
|
675
|
+
| GLM 5.3 Flash | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | Z.ai flash tier, 1.3M context² |
|
|
676
|
+
| MiMo V2.6 Pro | `xiaomi/mimo-v2.6-pro` | $0.435 | $0.87 | Xiaomi flagship, 1M context³ |
|
|
626
677
|
|
|
627
678
|
¹ Muse Spark 1.3 is age-gated by OpenRouter: calls return HTTP 403 (`VisualAIAuthError`) until the account completes the 18+ confirmation at [openrouter.ai/settings/preferences](https://openrouter.ai/settings/preferences). It also reasons by default — expect several hundred reasoning tokens per call even with no `reasoningEffort` set.
|
|
628
679
|
|
|
629
680
|
² GLM 5.3 Flash reasons by default — expect one to two hundred reasoning tokens per call even with no `reasoningEffort` set, billed at the output rate. OpenRouter's own context cap for it is 1,048,576 tokens (Z.ai lists 1,310,720) and its output ceiling is 131,072.
|
|
630
681
|
|
|
682
|
+
³ MiMo V2.6 Pro also reasons by default: about 390 reasoning tokens per call with no `reasoningEffort` set, and about 790 at `medium`. OpenRouter serves it from two fp8 hosts at the same price, Xiaomi (~36 tok/s) and DeepInfra (~4 tok/s), so per-call latency varies widely with the host it is routed to.
|
|
683
|
+
|
|
684
|
+
⁴ Qwen3.8 Max and Qwen3.7 Plus reason past the 4096-token default on a large share of calls (4,500–5,400 reasoning tokens on the long ones), so **they get a 32768-token output budget automatically** at every reasoning effort. At the default, about half their `ask()` calls truncated in live testing; with the larger budget every call completed. Passing `maxTokens` explicitly still wins, and a call that used the whole budget would cost at most about $0.20 on Qwen3.8 Max.
|
|
685
|
+
|
|
686
|
+
⁵ Kimi K2.7 Code fails roughly one call in ten by answering in prose instead of JSON. At the 4096 default those calls surface as `VisualAITruncationError`, and with a larger `maxTokens` they finish and throw `VisualAIResponseParseError` instead, so raising the budget does not help. Retry failed calls.
|
|
687
|
+
|
|
631
688
|
Meta also publishes `meta/muse-spark-1.3-contributor`, the same model at $0.10 / $0.20 per MTok — about 12x cheaper — because Meta uses everything submitted through it for product improvement. It has **no named constant** (`Model.OpenRouter` does not expose it) and never appears by default anywhere in this library, so using it takes a deliberate, explicit choice: pass the slug directly as a plain string, `visualAI({ model: "meta/muse-spark-1.3-contributor" })`. Any OpenRouter slug works this way — see the note above the table — and cost tracking works correctly once you opt in. OpenRouter itself blocks it with HTTP 404 (`paid-model-training-violation-by-account`) until the account's privacy settings allow training endpoints, at [openrouter.ai/settings/privacy](https://openrouter.ai/settings/privacy). Only use it if sending your screenshots to Meta for training is a trade you've deliberately made.
|
|
632
689
|
|
|
633
690
|
`qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
|