visual-ai-assertions 0.22.0 → 0.25.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +49 -9
- package/dist/index.cjs +393 -56
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +106 -17
- package/dist/index.d.ts +106 -17
- package/dist/index.js +393 -56
- package/dist/index.js.map +1 -1
- package/package.json +8 -8
package/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# visual-ai-assertions
|
|
2
2
|
|
|
3
|
-
AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, and
|
|
3
|
+
AI-powered visual assertions for E2E tests. Send screenshots — or short video recordings — to Claude, GPT, Gemini — or Grok, Kimi, Qwen, and GLM via OpenRouter — and get structured, typed results.
|
|
4
4
|
|
|
5
5
|
## Installation
|
|
6
6
|
|
|
@@ -11,7 +11,7 @@ npm install visual-ai-assertions
|
|
|
11
11
|
# Optional: install additional provider SDKs
|
|
12
12
|
npm install @anthropic-ai/sdk # for Claude
|
|
13
13
|
npm install @google/genai # for Gemini
|
|
14
|
-
# OpenRouter (Grok, Kimi, Qwen, ...) uses the OpenAI SDK — no extra install
|
|
14
|
+
# OpenRouter (Grok, Kimi, Qwen, GLM, ...) uses the OpenAI SDK — no extra install
|
|
15
15
|
|
|
16
16
|
# Zod is a peer dependency
|
|
17
17
|
npm install zod
|
|
@@ -352,16 +352,22 @@ Oversized images are automatically resized to provider limits.
|
|
|
352
352
|
|
|
353
353
|
`ai.check()` and `ai.ask()` also accept short video recordings (`.mp4`, `.webm`, `.mov`, `.mkv`) — useful for asserting on transient UI like toast messages. Accepted shapes are file path, `data:video/...;base64,...` URL, raw base64 string, `Buffer`, and `Uint8Array`. HTTP/HTTPS URLs are not supported for video inputs — fetch the bytes yourself first.
|
|
354
354
|
|
|
355
|
+
How the video reaches the model depends on the provider. **Google models receive the video itself** (native delivery): Gemini samples and tokenises it server-side and also hears the audio track. **Every other provider gets frames sampled with ffmpeg** and sent as an ordered image timeline. Both paths enforce the same duration cap before any provider call, and both give you per-statement timestamps.
|
|
356
|
+
|
|
355
357
|
```typescript
|
|
356
358
|
// Playwright recording on disk
|
|
357
359
|
const result = await ai.check("./trace/video/recording.webm", [
|
|
358
360
|
'A success toast with text "Saved" briefly appears',
|
|
359
361
|
]);
|
|
362
|
+
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
363
|
+
|
|
364
|
+
// Google model → the video was sent natively
|
|
365
|
+
console.log(result.video);
|
|
366
|
+
// { durationSeconds: 4.0, fps: 1, mimeType: "video/webm", delivery: "inline" }
|
|
360
367
|
|
|
361
|
-
//
|
|
368
|
+
// Any other provider → frames were sampled
|
|
362
369
|
console.log(result.frames);
|
|
363
|
-
// { count:
|
|
364
|
-
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
370
|
+
// { count: 3, timestampsSeconds: [0.5, 2.5, 3.5], durationSeconds: 4.0, droppedUnchanged: 1 }
|
|
365
371
|
|
|
366
372
|
// Override sampling — defaults are 1 fps, max 10 frames, max 10 s of video
|
|
367
373
|
await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
@@ -369,9 +375,37 @@ await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
|
369
375
|
});
|
|
370
376
|
```
|
|
371
377
|
|
|
372
|
-
`maxFrames` is hard-capped at 60 to keep memory bounded
|
|
378
|
+
`fps` applies to both paths: it is the ffmpeg sampling rate, and for native delivery it is passed to the provider as its sampling rate. `maxFrames` and `dedupe` apply to frame sampling only; `maxFrames` is hard-capped at 60 to keep memory bounded, and frames are downscaled so the longer edge fits within 1568 px before being sent.
|
|
379
|
+
|
|
380
|
+
**Choosing the delivery.** `video.mode` is `"auto"` by default: native where the provider supports it, frames elsewhere. Set `"frames"` to sample frames on a Google model too, or `"native"` to insist on native delivery — that throws `VisualAIConfigError` on providers without it.
|
|
381
|
+
|
|
382
|
+
```typescript
|
|
383
|
+
// Sample frames on Gemini instead of sending the video
|
|
384
|
+
await ai.check("./clip.mp4", ["Loader disappears"], { video: { mode: "frames" } });
|
|
385
|
+
|
|
386
|
+
// Fail loudly if the configured model cannot take video natively
|
|
387
|
+
await ai.ask("./clip.mp4", "What happens?", { video: { mode: "native" } });
|
|
388
|
+
```
|
|
389
|
+
|
|
390
|
+
Native delivery is worth preferring where available. Measured on a 15 s recording with a seven-statement `check()` on Gemini 3.8 Flash, native video at 1 fps cost about 60% of the frames path with dedupe on and answered about 1.5 s faster, with identical verdicts and correct timestamps on every run; Gemini tokenises a video frame at roughly 60 tokens against about 1,080 for the same frame sent as an image. Videos over 19 MB are uploaded through the Gemini Files API automatically (`delivery: "file"`) and deleted again afterwards. For `ask()`, native results carry `timestampReferences` (seconds) in place of the frame-indexed `frameReferences`.
|
|
373
391
|
|
|
374
|
-
|
|
392
|
+
**Unchanged frames are dropped.** A recording of a mostly static screen would otherwise pay input tokens for every near-identical sample, so by default each sampled frame is compared against the most recently kept frame and dropped when fewer than 0.1% of its pixels changed — roughly a 37x37 px region on a 1568x880 frame. A toast, a loader, or a dialog comfortably clears that; compression noise and a blinking text caret do not. The first frame is always kept, `durationSeconds` still covers the whole clip, and the prompt tells the model how many frames were dropped and that the screen stayed unchanged between the listed timestamps. `result.frames.droppedUnchanged` reports how many were dropped.
|
|
393
|
+
|
|
394
|
+
```typescript
|
|
395
|
+
// Keep every sampled frame
|
|
396
|
+
await ai.check("./clip.webm", ["A spinner is visible throughout"], {
|
|
397
|
+
video: { dedupe: false },
|
|
398
|
+
});
|
|
399
|
+
|
|
400
|
+
// Keep smaller changes (fraction of pixels that must change, in (0, 1])
|
|
401
|
+
await ai.check("./clip.webm", ["The 24 px status dot turns green"], {
|
|
402
|
+
video: { dedupe: { threshold: 0.0002 } },
|
|
403
|
+
});
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
A change smaller than the threshold — a lone 24x24 px icon on a full-size frame is about 0.04% — is treated as no change, so lower `threshold` or pass `dedupe: false` when the assertion is about something that small.
|
|
407
|
+
|
|
408
|
+
How it works: on the frames path the library samples frames with ffmpeg, drops the ones that did not visibly change, and sends the rest to the provider as an ordered timeline; on the native path it probes the duration, then hands the provider the video and a matching prompt. Either way a statement passes when it is true at any point, unless its wording specifies otherwise (e.g. "throughout"). Template helpers (`accessibility`, `layout`, `pageLoad`, `content`, `elementsVisible`, `elementsHidden`) are image-only — pass video to `check()` or `ask()` instead.
|
|
375
409
|
|
|
376
410
|
**ffmpeg setup.** Video support works out of the box — `fluent-ffmpeg`, `@ffmpeg-installer/ffmpeg`, and `@ffprobe-installer/ffprobe` ship as regular dependencies and bundle platform-specific ffmpeg/ffprobe binaries. If you ran `npm install` you already have everything you need. On platforms where the prebuilt binary is unavailable (or if you've pruned dependencies), `check()` and `ask()` throw `VisualAIVideoError` (import from `visual-ai-assertions` to `instanceof`-narrow it) when called with video input.
|
|
377
411
|
|
|
@@ -384,7 +418,7 @@ const frames = [await page.screenshot(), await page.screenshot()];
|
|
|
384
418
|
|
|
385
419
|
const result = await ai.check({ frames }, ['A success toast with text "Saved" appears']);
|
|
386
420
|
console.log(result.frames);
|
|
387
|
-
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1 }
|
|
421
|
+
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1, droppedUnchanged: 0 }
|
|
388
422
|
|
|
389
423
|
// Control timestamps: give an explicit fps, or per-frame timestampSeconds
|
|
390
424
|
await ai.ask(
|
|
@@ -398,7 +432,7 @@ await ai.ask(
|
|
|
398
432
|
);
|
|
399
433
|
```
|
|
400
434
|
|
|
401
|
-
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path.
|
|
435
|
+
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path — pre-sampled frames are always sent as frames, even to Google models. Frames that did not visibly change from the preceding kept frame are dropped exactly as for video input; control it with `dedupe` on the `FramesInput` itself, e.g. `{ frames, dedupe: false }` or `{ frames, dedupe: { threshold: 0.0002 } }`.
|
|
402
436
|
|
|
403
437
|
### Formatting & Assertion Helpers
|
|
404
438
|
|
|
@@ -511,7 +545,10 @@ import type {
|
|
|
511
545
|
CheckResult,
|
|
512
546
|
CompareResult,
|
|
513
547
|
Frame,
|
|
548
|
+
FrameDedupeOptions,
|
|
514
549
|
MediaInput,
|
|
550
|
+
NativeVideoMetadata,
|
|
551
|
+
VideoDeliveryMode,
|
|
515
552
|
SupportedMimeType,
|
|
516
553
|
SupportedVideoMimeType,
|
|
517
554
|
VideoFramesMetadata,
|
|
@@ -622,9 +659,12 @@ Any [OpenRouter](https://openrouter.ai/models) model slug (always `vendor/model`
|
|
|
622
659
|
| Qwen3.8 Max | `qwen/qwen3.8-max` | $2 | $6 | First Max tier with image input |
|
|
623
660
|
| Qwen3.7 Plus | `qwen/qwen3.7-plus` | $0.32 | $1.28 | Cost-effective, GUI/screen-reading |
|
|
624
661
|
| Qwen3.6 Flash | `qwen/qwen3.6-flash` | $0.19 | $1.13 | **Default** — cheap flash vision tier |
|
|
662
|
+
| GLM 5.3 Flash | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | Z.ai flash tier, 1.3M context² |
|
|
625
663
|
|
|
626
664
|
¹ Muse Spark 1.3 is age-gated by OpenRouter: calls return HTTP 403 (`VisualAIAuthError`) until the account completes the 18+ confirmation at [openrouter.ai/settings/preferences](https://openrouter.ai/settings/preferences). It also reasons by default — expect several hundred reasoning tokens per call even with no `reasoningEffort` set.
|
|
627
665
|
|
|
666
|
+
² GLM 5.3 Flash reasons by default — expect one to two hundred reasoning tokens per call even with no `reasoningEffort` set, billed at the output rate. OpenRouter's own context cap for it is 1,048,576 tokens (Z.ai lists 1,310,720) and its output ceiling is 131,072.
|
|
667
|
+
|
|
628
668
|
Meta also publishes `meta/muse-spark-1.3-contributor`, the same model at $0.10 / $0.20 per MTok — about 12x cheaper — because Meta uses everything submitted through it for product improvement. It has **no named constant** (`Model.OpenRouter` does not expose it) and never appears by default anywhere in this library, so using it takes a deliberate, explicit choice: pass the slug directly as a plain string, `visualAI({ model: "meta/muse-spark-1.3-contributor" })`. Any OpenRouter slug works this way — see the note above the table — and cost tracking works correctly once you opt in. OpenRouter itself blocks it with HTTP 404 (`paid-model-training-violation-by-account`) until the account's privacy settings allow training endpoints, at [openrouter.ai/settings/privacy](https://openrouter.ai/settings/privacy). Only use it if sending your screenshots to Meta for training is a trade you've deliberately made.
|
|
629
669
|
|
|
630
670
|
`qwen/qwen3.7-max` and the DeepSeek V4 family (`deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, and dated variants such as `deepseek/deepseek-v4-pro-0813`) are not listed because they accept no image input on OpenRouter.
|