@fortemate/vega-vvd-driver 0.2.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -12
- package/dist/record.js +83 -15
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -6,7 +6,7 @@ Drive the Vega Virtual Device from scripts and AI agents: press remote keys, tak
|
|
|
6
6
|
|
|
7
7
|
## Why
|
|
8
8
|
|
|
9
|
-
The Vega SDK's `vega` command installs and launches apps on the Vega Virtual Device (VVD), but has no command to press a remote key or take a screenshot. Amazon's [Appium Vega driver](https://developer.amazon.com/docs/vega/0.24/appium-install.html) does both, and much more: it runs UI tests that find elements, on the VVD and on a Fire TV Stick. It needs an Appium 2 server, the driver package and the device's automation toolkit switched on, with Node 22 or earlier.
|
|
9
|
+
The Vega SDK's `vega` command installs and launches apps on the Vega Virtual Device (VVD), but has no command to press a remote key or take a screenshot. On a Fire TV Stick, the device's own `gwsi-tool-screenshooter` takes one through `vega exec vda shell`, but the [VDA reference](https://developer.amazon.com/docs/vega/0.24/vda-tools.html) notes that the VVD doesn't support it. Amazon's [Appium Vega driver](https://developer.amazon.com/docs/vega/0.24/appium-install.html) does both, and much more: it runs UI tests that find elements, on the VVD and on a Fire TV Stick. It needs an Appium 2 server, the driver package and the device's automation toolkit switched on, with Node 22 or earlier.
|
|
10
10
|
|
|
11
11
|
This driver is for lighter jobs on the VVD: a key press or a screenshot from a shell script, a video with sound, every frame of an animation, and an AI coding agent that can see what it built. It talks to the Android emulator that the VVD is built on, through the emulator's own gRPC API and console, so there is nothing to install on the device and no server to run.
|
|
12
12
|
|
|
@@ -14,15 +14,17 @@ Its methods were worked out while building [Dice Chess for Fire TV](https://gith
|
|
|
14
14
|
|
|
15
15
|
## Appium or this driver?
|
|
16
16
|
|
|
17
|
-
| To… |
|
|
18
|
-
|
|
|
19
|
-
| find elements, read their text, run a test suite |
|
|
20
|
-
| test on a Fire TV Stick |
|
|
21
|
-
| press keys or take a screenshot from a shell script |
|
|
22
|
-
| record a video with sound |
|
|
23
|
-
| capture every frame of an animation |
|
|
24
|
-
| let an AI agent see and operate the VVD over MCP |
|
|
25
|
-
| check the TV safe area |
|
|
17
|
+
| To… | Appium | this driver |
|
|
18
|
+
| :-------------------------------------------------- | :----: | :---------: |
|
|
19
|
+
| find elements, read their text, run a test suite | ✅ | — |
|
|
20
|
+
| test on a Fire TV Stick | ✅ | — |
|
|
21
|
+
| press keys or take a screenshot from a shell script | ✅ | ✅ |
|
|
22
|
+
| record a video with sound | — | ✅ |
|
|
23
|
+
| capture every frame of an animation | — | ✅ |
|
|
24
|
+
| let an AI agent see and operate the VVD over MCP | — | ✅ |
|
|
25
|
+
| check the TV safe area | ◯ | ✅ |
|
|
26
|
+
|
|
27
|
+
✅ a good fit · ◯ possible with your own code (take a screenshot, then check its edges) · — not supported
|
|
26
28
|
|
|
27
29
|
## What you need
|
|
28
30
|
|
|
@@ -129,8 +131,8 @@ Measured on the VVD with Vega SDK 0.24.12112 on macOS:
|
|
|
129
131
|
- **Keys.** The emulator's gRPC `sendKey`, with Linux evdev codes, reaches apps: it is the path the VVD's own on-screen remote uses. OK is `KEY_KPENTER`, and apps receive it as `kpenter`, not the `select` the remote's documentation names. `KEY_SELECT` and `KEY_OK` never arrive, because the emulator's virtual keyboard does not declare them. Back is `KEY_BACK`; `KEY_ESC` does not reach an app as Back.
|
|
130
132
|
- **Dead ends.** The emulator console's `event send`, QEMU's `send-key` and `inputd-cli` on the device all report success and never reach an app.
|
|
131
133
|
- **Home cannot be sent.** `KEY_HOMEPAGE` (172), `KEY_F1` and 170, the code Amazon's Appium documentation gives for Home, all leave the app on screen. To get back to the launcher, run `vega device launch-app -d VirtualDevice -a com.amazon.keplerlauncherapp.main`.
|
|
132
|
-
- **Screenshots.** `getScreenshot` returns 1920x1080. A running process polls it in RGB at about 57 screenshots a second of a still screen, 17 ms each, and at 16 to 43 a second while the screen changes, 23 to 61 ms each (measured on 28 September 2026). `vvd screenshot` takes about half a second, most of it starting up. The emulator's `streamScreenshot` has delivered only its first frame while the screen kept changing, so the driver polls instead.
|
|
133
|
-
- **Audio.** `streamAudio` sends nothing while the device is silent. `record`
|
|
134
|
+
- **Screenshots.** `getScreenshot` returns 1920x1080. A running process polls it in RGB at about 57 screenshots a second of a still screen, 17 ms each, and at 16 to 43 a second while the screen changes, 23 to 61 ms each (measured on 28 September 2026). `vvd screenshot` takes about half a second, most of it starting up. The emulator's `streamScreenshot` has delivered only its first frame while the screen kept changing, so the driver polls instead. The device's `gwsi-tool-screenshooter` is for a Fire TV Stick: the VDA reference notes that the VVD doesn't support it, and on the VVD it wrote a 0-byte PNG or hung, with no error (28 September 2026).
|
|
135
|
+
- **Audio.** `streamAudio` sends nothing while the device is silent, and while a sound plays each packet's capture time wanders by about ±10 ms around where the packet before it ends. `record` lays the packets of one sound back to back and places the sound where their capture times agree best, then fills the gaps between sounds with silence. A sound stays in sync, and a long one, such as speech or music, plays without holes ([#13](https://github.com/fortemate/vega-vvd-driver/issues/13)). Measured on the VVD.
|
|
134
136
|
- **gRPC.** The endpoint is off after every start of the VVD, and the discovery file that the driver reads appears only once `grpc <port>` has been sent to the console. `vvd enable-grpc` does that.
|
|
135
137
|
|
|
136
138
|
## Troubleshooting
|
package/dist/record.js
CHANGED
|
@@ -7,8 +7,8 @@
|
|
|
7
7
|
// screen keeps changing. Polling does not stop: on the VVD a 1080p screenshot
|
|
8
8
|
// takes about 17 ms, or 23 to 61 ms while the screen changes. The audio comes
|
|
9
9
|
// from streamAudio. The emulator sends nothing while the device is silent, so
|
|
10
|
-
// the track is rebuilt on the video's clock from
|
|
11
|
-
// with silence in the gaps.
|
|
10
|
+
// the track is rebuilt on the video's clock from the packets' capture times,
|
|
11
|
+
// with silence in the gaps; see assembleAudio for how.
|
|
12
12
|
//
|
|
13
13
|
// However a recording ends, it cleans up after itself: ffmpeg is stopped, the
|
|
14
14
|
// audio stream is cancelled and the working directory is removed.
|
|
@@ -49,23 +49,91 @@ const abortable = (promise, signal) => {
|
|
|
49
49
|
};
|
|
50
50
|
const SAMPLE_RATE = 44100;
|
|
51
51
|
const FRAME_BYTES = 4; // 16-bit stereo
|
|
52
|
+
// A packet's capture time wanders by about ±10 ms around where the packet
|
|
53
|
+
// before it ends, and a stream starts with smaller packets (#13). Laid each at
|
|
54
|
+
// its own time, the packets of one sound would leave holes and cut into each
|
|
55
|
+
// other several times a second, which is heard as a rattle. So consecutive
|
|
56
|
+
// packets are laid back to back, the way the device played them, and each run
|
|
57
|
+
// of them is placed as a whole where the capture times agree best: their
|
|
58
|
+
// median. A run ends where the device sent nothing, a silence that leaves a
|
|
59
|
+
// step in time longer than SILENCE_US, and where it has drifted from the clock
|
|
60
|
+
// by more than DRIFT_US. Its first SETTLE_US are not held to the clock: there
|
|
61
|
+
// the small packets' capture times run ahead of their audio, by more than
|
|
62
|
+
// DRIFT_US at the start of some sounds, and splitting there would cut a hole
|
|
63
|
+
// just after the sound begins.
|
|
64
|
+
const SILENCE_US = 50_000;
|
|
65
|
+
const DRIFT_US = 100_000;
|
|
66
|
+
const SETTLE_US = 500_000;
|
|
67
|
+
const framesOf = (packet) => Math.floor(packet.pcm.length / FRAME_BYTES);
|
|
68
|
+
const usOf = (frames) => (frames / SAMPLE_RATE) * 1e6;
|
|
69
|
+
// The packets in the order they arrived, which is the order they were played,
|
|
70
|
+
// cut into runs at each silence.
|
|
71
|
+
const runsOf = (packets) => {
|
|
72
|
+
const runs = [];
|
|
73
|
+
let run = [];
|
|
74
|
+
for (const packet of packets) {
|
|
75
|
+
const last = run.at(-1);
|
|
76
|
+
if (last &&
|
|
77
|
+
Math.abs(packet.timestampUs - (last.timestampUs + usOf(framesOf(last)))) >
|
|
78
|
+
SILENCE_US) {
|
|
79
|
+
runs.push(run);
|
|
80
|
+
run = [];
|
|
81
|
+
}
|
|
82
|
+
run.push(packet);
|
|
83
|
+
}
|
|
84
|
+
if (run.length)
|
|
85
|
+
runs.push(run);
|
|
86
|
+
return runs;
|
|
87
|
+
};
|
|
88
|
+
// Where a run's first frame belongs on a clock that starts at `startUs`, in
|
|
89
|
+
// microseconds: each packet's capture time less the audio before it in the
|
|
90
|
+
// run, and the median of those, which jitter and the first small packets do
|
|
91
|
+
// not move.
|
|
92
|
+
const anchorOf = (run, startUs) => {
|
|
93
|
+
let before = 0;
|
|
94
|
+
const offsets = run.map((packet) => {
|
|
95
|
+
const offset = packet.timestampUs - startUs - usOf(before);
|
|
96
|
+
before += framesOf(packet);
|
|
97
|
+
return offset;
|
|
98
|
+
});
|
|
99
|
+
offsets.sort((a, b) => a - b);
|
|
100
|
+
return offsets[Math.floor(offsets.length / 2)];
|
|
101
|
+
};
|
|
102
|
+
// The run split where a packet, past the run's first SETTLE_US of audio, has
|
|
103
|
+
// drifted more than DRIFT_US from the place the run gives it, so a long sound
|
|
104
|
+
// follows the clock: the part before the first such packet, and the rest.
|
|
105
|
+
const steady = (run, startUs) => {
|
|
106
|
+
const anchor = anchorOf(run, startUs);
|
|
107
|
+
let before = 0;
|
|
108
|
+
for (let i = 0; i < run.length; i++) {
|
|
109
|
+
const drift = run[i].timestampUs - startUs - (anchor + usOf(before));
|
|
110
|
+
if (usOf(before) > SETTLE_US && Math.abs(drift) > DRIFT_US)
|
|
111
|
+
return [run.slice(0, i), run.slice(i)];
|
|
112
|
+
before += framesOf(run[i]);
|
|
113
|
+
}
|
|
114
|
+
return [[...run], []];
|
|
115
|
+
};
|
|
52
116
|
// Lays audio packets on a timeline that starts at `startUs` and lasts
|
|
53
|
-
// `seconds`, as 16-bit stereo PCM
|
|
54
|
-
//
|
|
55
|
-
//
|
|
117
|
+
// `seconds`, as 16-bit stereo PCM, with silence where nothing arrived. Audio
|
|
118
|
+
// placed before the start is cut, and past the end is dropped; the result is
|
|
119
|
+
// exactly as long as asked. Where two runs overlap, the later one is heard.
|
|
56
120
|
export const assembleAudio = (packets, startUs, seconds) => {
|
|
57
121
|
const track = Buffer.alloc(Math.round(seconds * SAMPLE_RATE) * FRAME_BYTES);
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
122
|
+
const pending = runsOf(packets);
|
|
123
|
+
while (pending.length) {
|
|
124
|
+
const [run, rest] = steady(pending.shift(), startUs);
|
|
125
|
+
if (rest.length)
|
|
126
|
+
pending.unshift(rest);
|
|
127
|
+
let offset = Math.round((anchorOf(run, startUs) / 1e6) * SAMPLE_RATE);
|
|
128
|
+
for (const packet of run) {
|
|
129
|
+
const frames = framesOf(packet);
|
|
130
|
+
const from = Math.max(0, -offset);
|
|
131
|
+
const at = offset + from;
|
|
132
|
+
offset += frames;
|
|
133
|
+
if (from >= frames || at * FRAME_BYTES >= track.length)
|
|
134
|
+
continue;
|
|
135
|
+
packet.pcm.copy(track, at * FRAME_BYTES, from * FRAME_BYTES, frames * FRAME_BYTES);
|
|
65
136
|
}
|
|
66
|
-
if (from >= frames || offset * FRAME_BYTES >= track.length)
|
|
67
|
-
continue;
|
|
68
|
-
packet.pcm.copy(track, offset * FRAME_BYTES, from * FRAME_BYTES, frames * FRAME_BYTES);
|
|
69
137
|
}
|
|
70
138
|
return track;
|
|
71
139
|
};
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@fortemate/vega-vvd-driver",
|
|
3
|
-
"version": "0.2.
|
|
3
|
+
"version": "0.2.1",
|
|
4
4
|
"description": "Unofficial driver for the Vega Virtual Device: remote keys, screenshots and video with sound, from scripts and AI agents. CLI, Node library and MCP server.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"vega",
|