visual-ai-assertions 0.23.0 → 0.25.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +44 -7
- package/dist/index.cjs +385 -55
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +105 -17
- package/dist/index.d.ts +105 -17
- package/dist/index.js +385 -55
- package/dist/index.js.map +1 -1
- package/package.json +8 -8
package/README.md
CHANGED
|
@@ -352,16 +352,22 @@ Oversized images are automatically resized to provider limits.
|
|
|
352
352
|
|
|
353
353
|
`ai.check()` and `ai.ask()` also accept short video recordings (`.mp4`, `.webm`, `.mov`, `.mkv`) — useful for asserting on transient UI like toast messages. Accepted shapes are file path, `data:video/...;base64,...` URL, raw base64 string, `Buffer`, and `Uint8Array`. HTTP/HTTPS URLs are not supported for video inputs — fetch the bytes yourself first.
|
|
354
354
|
|
|
355
|
+
How the video reaches the model depends on the provider. **Google models receive the video itself** (native delivery): Gemini samples and tokenises it server-side and also hears the audio track. **Every other provider gets frames sampled with ffmpeg** and sent as an ordered image timeline. Both paths enforce the same duration cap before any provider call, and both give you per-statement timestamps.
|
|
356
|
+
|
|
355
357
|
```typescript
|
|
356
358
|
// Playwright recording on disk
|
|
357
359
|
const result = await ai.check("./trace/video/recording.webm", [
|
|
358
360
|
'A success toast with text "Saved" briefly appears',
|
|
359
361
|
]);
|
|
362
|
+
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
363
|
+
|
|
364
|
+
// Google model → the video was sent natively
|
|
365
|
+
console.log(result.video);
|
|
366
|
+
// { durationSeconds: 4.0, fps: 1, mimeType: "video/webm", delivery: "inline" }
|
|
360
367
|
|
|
361
|
-
//
|
|
368
|
+
// Any other provider → frames were sampled
|
|
362
369
|
console.log(result.frames);
|
|
363
|
-
// { count:
|
|
364
|
-
console.log(result.statements[0].timestampSeconds); // 3.5
|
|
370
|
+
// { count: 3, timestampsSeconds: [0.5, 2.5, 3.5], durationSeconds: 4.0, droppedUnchanged: 1 }
|
|
365
371
|
|
|
366
372
|
// Override sampling — defaults are 1 fps, max 10 frames, max 10 s of video
|
|
367
373
|
await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
@@ -369,9 +375,37 @@ await ai.check("./long-clip.mp4", ["Loader disappears"], {
|
|
|
369
375
|
});
|
|
370
376
|
```
|
|
371
377
|
|
|
372
|
-
`maxFrames` is hard-capped at 60 to keep memory bounded
|
|
378
|
+
`fps` applies to both paths: it is the ffmpeg sampling rate, and for native delivery it is passed to the provider as its sampling rate. `maxFrames` and `dedupe` apply to frame sampling only; `maxFrames` is hard-capped at 60 to keep memory bounded, and frames are downscaled so the longer edge fits within 1568 px before being sent.
|
|
379
|
+
|
|
380
|
+
**Choosing the delivery.** `video.mode` is `"auto"` by default: native where the provider supports it, frames elsewhere. Set `"frames"` to sample frames on a Google model too, or `"native"` to insist on native delivery — that throws `VisualAIConfigError` on providers without it.
|
|
381
|
+
|
|
382
|
+
```typescript
|
|
383
|
+
// Sample frames on Gemini instead of sending the video
|
|
384
|
+
await ai.check("./clip.mp4", ["Loader disappears"], { video: { mode: "frames" } });
|
|
385
|
+
|
|
386
|
+
// Fail loudly if the configured model cannot take video natively
|
|
387
|
+
await ai.ask("./clip.mp4", "What happens?", { video: { mode: "native" } });
|
|
388
|
+
```
|
|
389
|
+
|
|
390
|
+
Native delivery is worth preferring where available. Measured on a 15 s recording with a seven-statement `check()` on Gemini 3.8 Flash, native video at 1 fps cost about 60% of the frames path with dedupe on and answered about 1.5 s faster, with identical verdicts and correct timestamps on every run; Gemini tokenises a video frame at roughly 60 tokens against about 1,080 for the same frame sent as an image. Videos over 19 MB are uploaded through the Gemini Files API automatically (`delivery: "file"`) and deleted again afterwards. For `ask()`, native results carry `timestampReferences` (seconds) in place of the frame-indexed `frameReferences`.
|
|
391
|
+
|
|
392
|
+
**Unchanged frames are dropped.** A recording of a mostly static screen would otherwise pay input tokens for every near-identical sample, so by default each sampled frame is compared against the most recently kept frame and dropped when fewer than 0.1% of its pixels changed — roughly a 37x37 px region on a 1568x880 frame. A toast, a loader, or a dialog comfortably clears that; compression noise and a blinking text caret do not. The first frame is always kept, `durationSeconds` still covers the whole clip, and the prompt tells the model how many frames were dropped and that the screen stayed unchanged between the listed timestamps. `result.frames.droppedUnchanged` reports how many were dropped.
|
|
393
|
+
|
|
394
|
+
```typescript
|
|
395
|
+
// Keep every sampled frame
|
|
396
|
+
await ai.check("./clip.webm", ["A spinner is visible throughout"], {
|
|
397
|
+
video: { dedupe: false },
|
|
398
|
+
});
|
|
399
|
+
|
|
400
|
+
// Keep smaller changes (fraction of pixels that must change, in (0, 1])
|
|
401
|
+
await ai.check("./clip.webm", ["The 24 px status dot turns green"], {
|
|
402
|
+
video: { dedupe: { threshold: 0.0002 } },
|
|
403
|
+
});
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
A change smaller than the threshold — a lone 24x24 px icon on a full-size frame is about 0.04% — is treated as no change, so lower `threshold` or pass `dedupe: false` when the assertion is about something that small.
|
|
373
407
|
|
|
374
|
-
How it works: the library samples frames with ffmpeg and sends
|
|
408
|
+
How it works: on the frames path the library samples frames with ffmpeg, drops the ones that did not visibly change, and sends the rest to the provider as an ordered timeline; on the native path it probes the duration, then hands the provider the video and a matching prompt. Either way a statement passes when it is true at any point, unless its wording specifies otherwise (e.g. "throughout"). Template helpers (`accessibility`, `layout`, `pageLoad`, `content`, `elementsVisible`, `elementsHidden`) are image-only — pass video to `check()` or `ask()` instead.
|
|
375
409
|
|
|
376
410
|
**ffmpeg setup.** Video support works out of the box — `fluent-ffmpeg`, `@ffmpeg-installer/ffmpeg`, and `@ffprobe-installer/ffprobe` ship as regular dependencies and bundle platform-specific ffmpeg/ffprobe binaries. If you ran `npm install` you already have everything you need. On platforms where the prebuilt binary is unavailable (or if you've pruned dependencies), `check()` and `ask()` throw `VisualAIVideoError` (import from `visual-ai-assertions` to `instanceof`-narrow it) when called with video input.
|
|
377
411
|
|
|
@@ -384,7 +418,7 @@ const frames = [await page.screenshot(), await page.screenshot()];
|
|
|
384
418
|
|
|
385
419
|
const result = await ai.check({ frames }, ['A success toast with text "Saved" appears']);
|
|
386
420
|
console.log(result.frames);
|
|
387
|
-
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1 }
|
|
421
|
+
// { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1, droppedUnchanged: 0 }
|
|
388
422
|
|
|
389
423
|
// Control timestamps: give an explicit fps, or per-frame timestampSeconds
|
|
390
424
|
await ai.ask(
|
|
@@ -398,7 +432,7 @@ await ai.ask(
|
|
|
398
432
|
);
|
|
399
433
|
```
|
|
400
434
|
|
|
401
|
-
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path.
|
|
435
|
+
Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path — pre-sampled frames are always sent as frames, even to Google models. Frames that did not visibly change from the preceding kept frame are dropped exactly as for video input; control it with `dedupe` on the `FramesInput` itself, e.g. `{ frames, dedupe: false }` or `{ frames, dedupe: { threshold: 0.0002 } }`.
|
|
402
436
|
|
|
403
437
|
### Formatting & Assertion Helpers
|
|
404
438
|
|
|
@@ -511,7 +545,10 @@ import type {
|
|
|
511
545
|
CheckResult,
|
|
512
546
|
CompareResult,
|
|
513
547
|
Frame,
|
|
548
|
+
FrameDedupeOptions,
|
|
514
549
|
MediaInput,
|
|
550
|
+
NativeVideoMetadata,
|
|
551
|
+
VideoDeliveryMode,
|
|
515
552
|
SupportedMimeType,
|
|
516
553
|
SupportedVideoMimeType,
|
|
517
554
|
VideoFramesMetadata,
|