visual-ai-assertions 0.23.0 → 0.25.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -352,16 +352,22 @@ Oversized images are automatically resized to provider limits.
352
352
 
353
353
  `ai.check()` and `ai.ask()` also accept short video recordings (`.mp4`, `.webm`, `.mov`, `.mkv`) — useful for asserting on transient UI like toast messages. Accepted shapes are file path, `data:video/...;base64,...` URL, raw base64 string, `Buffer`, and `Uint8Array`. HTTP/HTTPS URLs are not supported for video inputs — fetch the bytes yourself first.
354
354
 
355
+ How the video reaches the model depends on the provider. **Google models receive the video itself** (native delivery): Gemini samples and tokenises it server-side and also hears the audio track. **Every other provider gets frames sampled with ffmpeg** and sent as an ordered image timeline. Both paths enforce the same duration cap before any provider call, and both give you per-statement timestamps.
356
+
355
357
  ```typescript
356
358
  // Playwright recording on disk
357
359
  const result = await ai.check("./trace/video/recording.webm", [
358
360
  'A success toast with text "Saved" briefly appears',
359
361
  ]);
362
+ console.log(result.statements[0].timestampSeconds); // 3.5
363
+
364
+ // Google model → the video was sent natively
365
+ console.log(result.video);
366
+ // { durationSeconds: 4.0, fps: 1, mimeType: "video/webm", delivery: "inline" }
360
367
 
361
- // Result includes frame metadata + per-statement timestamps
368
+ // Any other provider → frames were sampled
362
369
  console.log(result.frames);
363
- // { count: 4, timestampsSeconds: [0.5, 1.5, 2.5, 3.5], durationSeconds: 4.0 }
364
- console.log(result.statements[0].timestampSeconds); // 3.5
370
+ // { count: 3, timestampsSeconds: [0.5, 2.5, 3.5], durationSeconds: 4.0, droppedUnchanged: 1 }
365
371
 
366
372
  // Override sampling — defaults are 1 fps, max 10 frames, max 10 s of video
367
373
  await ai.check("./long-clip.mp4", ["Loader disappears"], {
@@ -369,9 +375,37 @@ await ai.check("./long-clip.mp4", ["Loader disappears"], {
369
375
  });
370
376
  ```
371
377
 
372
- `maxFrames` is hard-capped at 60 to keep memory bounded. Frames are downscaled so the longer edge fits within 1568 px before being sent to the provider.
378
+ `fps` applies to both paths: it is the ffmpeg sampling rate, and for native delivery it is passed to the provider as its sampling rate. `maxFrames` and `dedupe` apply to frame sampling only; `maxFrames` is hard-capped at 60 to keep memory bounded, and frames are downscaled so the longer edge fits within 1568 px before being sent.
379
+
380
+ **Choosing the delivery.** `video.mode` is `"auto"` by default: native where the provider supports it, frames elsewhere. Set `"frames"` to sample frames on a Google model too, or `"native"` to insist on native delivery — that throws `VisualAIConfigError` on providers without it.
381
+
382
+ ```typescript
383
+ // Sample frames on Gemini instead of sending the video
384
+ await ai.check("./clip.mp4", ["Loader disappears"], { video: { mode: "frames" } });
385
+
386
+ // Fail loudly if the configured model cannot take video natively
387
+ await ai.ask("./clip.mp4", "What happens?", { video: { mode: "native" } });
388
+ ```
389
+
390
+ Native delivery is worth preferring where available. Measured on a 15 s recording with a seven-statement `check()` on Gemini 3.8 Flash, native video at 1 fps cost about 60% of the frames path with dedupe on and answered about 1.5 s faster, with identical verdicts and correct timestamps on every run; Gemini tokenises a video frame at roughly 60 tokens against about 1,080 for the same frame sent as an image. Videos over 19 MB are uploaded through the Gemini Files API automatically (`delivery: "file"`) and deleted again afterwards. For `ask()`, native results carry `timestampReferences` (seconds) in place of the frame-indexed `frameReferences`.
391
+
392
+ **Unchanged frames are dropped.** A recording of a mostly static screen would otherwise pay input tokens for every near-identical sample, so by default each sampled frame is compared against the most recently kept frame and dropped when fewer than 0.1% of its pixels changed — roughly a 37x37 px region on a 1568x880 frame. A toast, a loader, or a dialog comfortably clears that; compression noise and a blinking text caret do not. The first frame is always kept, `durationSeconds` still covers the whole clip, and the prompt tells the model how many frames were dropped and that the screen stayed unchanged between the listed timestamps. `result.frames.droppedUnchanged` reports how many were dropped.
393
+
394
+ ```typescript
395
+ // Keep every sampled frame
396
+ await ai.check("./clip.webm", ["A spinner is visible throughout"], {
397
+ video: { dedupe: false },
398
+ });
399
+
400
+ // Keep smaller changes (fraction of pixels that must change, in (0, 1])
401
+ await ai.check("./clip.webm", ["The 24 px status dot turns green"], {
402
+ video: { dedupe: { threshold: 0.0002 } },
403
+ });
404
+ ```
405
+
406
+ A change smaller than the threshold — a lone 24x24 px icon on a full-size frame is about 0.04% — is treated as no change, so lower `threshold` or pass `dedupe: false` when the assertion is about something that small.
373
407
 
374
- How it works: the library samples frames with ffmpeg and sends them to the provider as an ordered timeline. A statement passes when it is true at any sampled frame, unless its wording specifies otherwise (e.g. "throughout"). Template helpers (`accessibility`, `layout`, `pageLoad`, `content`, `elementsVisible`, `elementsHidden`) are image-only — pass video to `check()` or `ask()` instead.
408
+ How it works: on the frames path the library samples frames with ffmpeg, drops the ones that did not visibly change, and sends the rest to the provider as an ordered timeline; on the native path it probes the duration, then hands the provider the video and a matching prompt. Either way a statement passes when it is true at any point, unless its wording specifies otherwise (e.g. "throughout"). Template helpers (`accessibility`, `layout`, `pageLoad`, `content`, `elementsVisible`, `elementsHidden`) are image-only — pass video to `check()` or `ask()` instead.
375
409
 
376
410
  **ffmpeg setup.** Video support works out of the box — `fluent-ffmpeg`, `@ffmpeg-installer/ffmpeg`, and `@ffprobe-installer/ffprobe` ship as regular dependencies and bundle platform-specific ffmpeg/ffprobe binaries. If you ran `npm install` you already have everything you need. On platforms where the prebuilt binary is unavailable (or if you've pruned dependencies), `check()` and `ask()` throw `VisualAIVideoError` (import from `visual-ai-assertions` to `instanceof`-narrow it) when called with video input.
377
411
 
@@ -384,7 +418,7 @@ const frames = [await page.screenshot(), await page.screenshot()];
384
418
 
385
419
  const result = await ai.check({ frames }, ['A success toast with text "Saved" appears']);
386
420
  console.log(result.frames);
387
- // { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1 }
421
+ // { count: 2, timestampsSeconds: [0, 1], durationSeconds: 1, droppedUnchanged: 0 }
388
422
 
389
423
  // Control timestamps: give an explicit fps, or per-frame timestampSeconds
390
424
  await ai.ask(
@@ -398,7 +432,7 @@ await ai.ask(
398
432
  );
399
433
  ```
400
434
 
401
- Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path.
435
+ Timestamps for bare frames are derived as `index / fps` (default `fps` is `1`); a per-frame `timestampSeconds` overrides that. The frame count is subject to the same 60-frame hard cap as video sampling, and the `video` sampling option is ignored for this path — pre-sampled frames are always sent as frames, even to Google models. Frames that did not visibly change from the preceding kept frame are dropped exactly as for video input; control it with `dedupe` on the `FramesInput` itself, e.g. `{ frames, dedupe: false }` or `{ frames, dedupe: { threshold: 0.0002 } }`.
402
436
 
403
437
  ### Formatting & Assertion Helpers
404
438
 
@@ -511,7 +545,10 @@ import type {
511
545
  CheckResult,
512
546
  CompareResult,
513
547
  Frame,
548
+ FrameDedupeOptions,
514
549
  MediaInput,
550
+ NativeVideoMetadata,
551
+ VideoDeliveryMode,
515
552
  SupportedMimeType,
516
553
  SupportedVideoMimeType,
517
554
  VideoFramesMetadata,