@fortemate/vega-vvd-driver 0.2.0 → 0.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +16 -12
- package/dist/mcp.js +25 -4
- package/dist/record.js +83 -15
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -6,7 +6,7 @@ Drive the Vega Virtual Device from scripts and AI agents: press remote keys, tak
|
|
|
6
6
|
|
|
7
7
|
## Why
|
|
8
8
|
|
|
9
|
-
The Vega SDK's `vega` command installs and launches apps on the Vega Virtual Device (VVD), but has no command to press a remote key or take a screenshot. Amazon's [Appium Vega driver](https://developer.amazon.com/docs/vega/0.24/appium-install.html) does both, and much more: it runs UI tests that find elements, on the VVD and on a Fire TV Stick. It needs an Appium 2 server, the driver package and the device's automation toolkit switched on, with Node 22 or earlier.
|
|
9
|
+
The Vega SDK's `vega` command installs and launches apps on the Vega Virtual Device (VVD), but has no command to press a remote key or take a screenshot. On a Fire TV Stick, the device's own `gwsi-tool-screenshooter` takes one through `vega exec vda shell`, but the [VDA reference](https://developer.amazon.com/docs/vega/0.24/vda-tools.html) notes that the VVD doesn't support it. Amazon's [Appium Vega driver](https://developer.amazon.com/docs/vega/0.24/appium-install.html) does both, and much more: it runs UI tests that find elements, on the VVD and on a Fire TV Stick. It needs an Appium 2 server, the driver package and the device's automation toolkit switched on, with Node 22 or earlier.
|
|
10
10
|
|
|
11
11
|
This driver is for lighter jobs on the VVD: a key press or a screenshot from a shell script, a video with sound, every frame of an animation, and an AI coding agent that can see what it built. It talks to the Android emulator that the VVD is built on, through the emulator's own gRPC API and console, so there is nothing to install on the device and no server to run.
|
|
12
12
|
|
|
@@ -14,15 +14,17 @@ Its methods were worked out while building [Dice Chess for Fire TV](https://gith
|
|
|
14
14
|
|
|
15
15
|
## Appium or this driver?
|
|
16
16
|
|
|
17
|
-
| To… |
|
|
18
|
-
|
|
|
19
|
-
| find elements, read their text, run a test suite |
|
|
20
|
-
| test on a Fire TV Stick |
|
|
21
|
-
| press keys or take a screenshot from a shell script |
|
|
22
|
-
| record a video with sound |
|
|
23
|
-
| capture every frame of an animation |
|
|
24
|
-
| let an AI agent see and operate the VVD over MCP |
|
|
25
|
-
| check the TV safe area |
|
|
17
|
+
| To… | Appium | this driver |
|
|
18
|
+
| :-------------------------------------------------- | :----: | :---------: |
|
|
19
|
+
| find elements, read their text, run a test suite | ✅ | — |
|
|
20
|
+
| test on a Fire TV Stick | ✅ | — |
|
|
21
|
+
| press keys or take a screenshot from a shell script | ✅ | ✅ |
|
|
22
|
+
| record a video with sound | — | ✅ |
|
|
23
|
+
| capture every frame of an animation | — | ✅ |
|
|
24
|
+
| let an AI agent see and operate the VVD over MCP | — | ✅ |
|
|
25
|
+
| check the TV safe area | ◯ | ✅ |
|
|
26
|
+
|
|
27
|
+
✅ a good fit · ◯ possible with your own code (take a screenshot, then check its edges) · — not supported
|
|
26
28
|
|
|
27
29
|
## What you need
|
|
28
30
|
|
|
@@ -90,6 +92,8 @@ Then ask, for example: "Open Settings in the app on the Virtual Device, turn the
|
|
|
90
92
|
|
|
91
93
|
A model chooses the arguments, so they are bounded: at most 100 key presses per call, recordings of up to 600 seconds, and `record_video` writes only a new `.mp4`, `.mov` or `.mkv` file, unless it is told to `overwrite` one. When the client cancels a call, its presses, wait or recording stop. `vvd mcp --pid <n>` ties the server to one device.
|
|
92
94
|
|
|
95
|
+
Each tool declares MCP annotations, for a client that asks before a call that changes something. Four tools only read. `enable_grpc` only turns on the emulator's gRPC endpoint. `press_keys` and `record_video` are marked destructive: a key press can confirm a delete in the app on screen, and `overwrite` replaces a file. None reaches beyond the local device.
|
|
96
|
+
|
|
93
97
|
## As a library
|
|
94
98
|
|
|
95
99
|
```ts
|
|
@@ -129,8 +133,8 @@ Measured on the VVD with Vega SDK 0.24.12112 on macOS:
|
|
|
129
133
|
- **Keys.** The emulator's gRPC `sendKey`, with Linux evdev codes, reaches apps: it is the path the VVD's own on-screen remote uses. OK is `KEY_KPENTER`, and apps receive it as `kpenter`, not the `select` the remote's documentation names. `KEY_SELECT` and `KEY_OK` never arrive, because the emulator's virtual keyboard does not declare them. Back is `KEY_BACK`; `KEY_ESC` does not reach an app as Back.
|
|
130
134
|
- **Dead ends.** The emulator console's `event send`, QEMU's `send-key` and `inputd-cli` on the device all report success and never reach an app.
|
|
131
135
|
- **Home cannot be sent.** `KEY_HOMEPAGE` (172), `KEY_F1` and 170, the code Amazon's Appium documentation gives for Home, all leave the app on screen. To get back to the launcher, run `vega device launch-app -d VirtualDevice -a com.amazon.keplerlauncherapp.main`.
|
|
132
|
-
- **Screenshots.** `getScreenshot` returns 1920x1080. A running process polls it in RGB at about 57 screenshots a second of a still screen, 17 ms each, and at 16 to 43 a second while the screen changes, 23 to 61 ms each (measured on 28 September 2026). `vvd screenshot` takes about half a second, most of it starting up. The emulator's `streamScreenshot` has delivered only its first frame while the screen kept changing, so the driver polls instead.
|
|
133
|
-
- **Audio.** `streamAudio` sends nothing while the device is silent. `record`
|
|
136
|
+
- **Screenshots.** `getScreenshot` returns 1920x1080. A running process polls it in RGB at about 57 screenshots a second of a still screen, 17 ms each, and at 16 to 43 a second while the screen changes, 23 to 61 ms each (measured on 28 September 2026). `vvd screenshot` takes about half a second, most of it starting up. The emulator's `streamScreenshot` has delivered only its first frame while the screen kept changing, so the driver polls instead. The device's `gwsi-tool-screenshooter` is for a Fire TV Stick: the VDA reference notes that the VVD doesn't support it, and on the VVD it wrote a 0-byte PNG or hung, with no error (28 September 2026).
|
|
137
|
+
- **Audio.** `streamAudio` sends nothing while the device is silent, and while a sound plays each packet's capture time wanders by about ±10 ms around where the packet before it ends. `record` lays the packets of one sound back to back and places the sound where their capture times agree best, then fills the gaps between sounds with silence. A sound stays in sync, and a long one, such as speech or music, plays without holes ([#13](https://github.com/fortemate/vega-vvd-driver/issues/13)). Measured on the VVD.
|
|
134
138
|
- **gRPC.** The endpoint is off after every start of the VVD, and the discovery file that the driver reads appears only once `grpc <port>` has been sent to the console. `vvd enable-grpc` does that.
|
|
135
139
|
|
|
136
140
|
## Troubleshooting
|
package/dist/mcp.js
CHANGED
|
@@ -71,7 +71,7 @@ export const createServer = (options = {}) => {
|
|
|
71
71
|
server.registerTool('list_devices', {
|
|
72
72
|
title: 'List running Vega Virtual Devices',
|
|
73
73
|
description: 'Lists the running Vega Virtual Devices whose gRPC endpoint is on. An empty list usually means gRPC is off: call enable_grpc.',
|
|
74
|
-
annotations: { readOnlyHint: true },
|
|
74
|
+
annotations: { readOnlyHint: true, openWorldHint: false },
|
|
75
75
|
}, async () => {
|
|
76
76
|
const devices = findEmulators({ directories: options.directories }).map(({ pid, grpcPort, consolePort, avdName }) => ({
|
|
77
77
|
pid,
|
|
@@ -101,6 +101,12 @@ export const createServer = (options = {}) => {
|
|
|
101
101
|
.optional()
|
|
102
102
|
.describe('The emulator console, an even port; default 5554'),
|
|
103
103
|
},
|
|
104
|
+
annotations: {
|
|
105
|
+
readOnlyHint: false,
|
|
106
|
+
destructiveHint: false,
|
|
107
|
+
idempotentHint: true,
|
|
108
|
+
openWorldHint: false,
|
|
109
|
+
},
|
|
104
110
|
}, async ({ grpc_port, console_port }) => {
|
|
105
111
|
try {
|
|
106
112
|
await enableGrpc(grpc_port ?? 8554, { port: console_port });
|
|
@@ -136,6 +142,14 @@ export const createServer = (options = {}) => {
|
|
|
136
142
|
.optional()
|
|
137
143
|
.describe('Return a screenshot once the keys are pressed'),
|
|
138
144
|
},
|
|
145
|
+
// A key press can confirm anything the app on screen offers, a delete
|
|
146
|
+
// included.
|
|
147
|
+
annotations: {
|
|
148
|
+
readOnlyHint: false,
|
|
149
|
+
destructiveHint: true,
|
|
150
|
+
idempotentHint: false,
|
|
151
|
+
openWorldHint: false,
|
|
152
|
+
},
|
|
139
153
|
}, async ({ keys, gap_ms, screenshot_after }, { signal }) => {
|
|
140
154
|
try {
|
|
141
155
|
const steps = parseKeys(keys); // a typo fails before any device is touched
|
|
@@ -155,7 +169,7 @@ export const createServer = (options = {}) => {
|
|
|
155
169
|
server.registerTool('screenshot', {
|
|
156
170
|
title: 'Take a screenshot',
|
|
157
171
|
description: 'Returns the current screen of the Vega Virtual Device as a PNG image (1920x1080).',
|
|
158
|
-
annotations: { readOnlyHint: true },
|
|
172
|
+
annotations: { readOnlyHint: true, openWorldHint: false },
|
|
159
173
|
}, async ({ signal }) => {
|
|
160
174
|
try {
|
|
161
175
|
return { content: [await screen(connected(), signal)] };
|
|
@@ -176,7 +190,7 @@ export const createServer = (options = {}) => {
|
|
|
176
190
|
.optional()
|
|
177
191
|
.describe('Default 5000'),
|
|
178
192
|
},
|
|
179
|
-
annotations: { readOnlyHint: true },
|
|
193
|
+
annotations: { readOnlyHint: true, openWorldHint: false },
|
|
180
194
|
}, async ({ timeout_ms }, { signal }) => {
|
|
181
195
|
try {
|
|
182
196
|
const d = connected();
|
|
@@ -211,6 +225,13 @@ export const createServer = (options = {}) => {
|
|
|
211
225
|
.optional()
|
|
212
226
|
.describe('Replace an existing file; default false'),
|
|
213
227
|
},
|
|
228
|
+
// Destructive only with overwrite, which replaces an existing file.
|
|
229
|
+
annotations: {
|
|
230
|
+
readOnlyHint: false,
|
|
231
|
+
destructiveHint: true,
|
|
232
|
+
idempotentHint: false,
|
|
233
|
+
openWorldHint: false,
|
|
234
|
+
},
|
|
214
235
|
}, async ({ file, seconds, fps, audio, overwrite }, { signal }) => {
|
|
215
236
|
try {
|
|
216
237
|
const path = videoPath(file, overwrite);
|
|
@@ -243,7 +264,7 @@ export const createServer = (options = {}) => {
|
|
|
243
264
|
.optional()
|
|
244
265
|
.describe('Fraction of each edge, default 0.05'),
|
|
245
266
|
},
|
|
246
|
-
annotations: { readOnlyHint: true },
|
|
267
|
+
annotations: { readOnlyHint: true, openWorldHint: false },
|
|
247
268
|
}, async ({ background, margin }, { signal }) => {
|
|
248
269
|
try {
|
|
249
270
|
const frame = await connected().screenshot('rgb', { signal });
|
package/dist/record.js
CHANGED
|
@@ -7,8 +7,8 @@
|
|
|
7
7
|
// screen keeps changing. Polling does not stop: on the VVD a 1080p screenshot
|
|
8
8
|
// takes about 17 ms, or 23 to 61 ms while the screen changes. The audio comes
|
|
9
9
|
// from streamAudio. The emulator sends nothing while the device is silent, so
|
|
10
|
-
// the track is rebuilt on the video's clock from
|
|
11
|
-
// with silence in the gaps.
|
|
10
|
+
// the track is rebuilt on the video's clock from the packets' capture times,
|
|
11
|
+
// with silence in the gaps; see assembleAudio for how.
|
|
12
12
|
//
|
|
13
13
|
// However a recording ends, it cleans up after itself: ffmpeg is stopped, the
|
|
14
14
|
// audio stream is cancelled and the working directory is removed.
|
|
@@ -49,23 +49,91 @@ const abortable = (promise, signal) => {
|
|
|
49
49
|
};
|
|
50
50
|
const SAMPLE_RATE = 44100;
|
|
51
51
|
const FRAME_BYTES = 4; // 16-bit stereo
|
|
52
|
+
// A packet's capture time wanders by about ±10 ms around where the packet
|
|
53
|
+
// before it ends, and a stream starts with smaller packets (#13). Laid each at
|
|
54
|
+
// its own time, the packets of one sound would leave holes and cut into each
|
|
55
|
+
// other several times a second, which is heard as a rattle. So consecutive
|
|
56
|
+
// packets are laid back to back, the way the device played them, and each run
|
|
57
|
+
// of them is placed as a whole where the capture times agree best: their
|
|
58
|
+
// median. A run ends where the device sent nothing, a silence that leaves a
|
|
59
|
+
// step in time longer than SILENCE_US, and where it has drifted from the clock
|
|
60
|
+
// by more than DRIFT_US. Its first SETTLE_US are not held to the clock: there
|
|
61
|
+
// the small packets' capture times run ahead of their audio, by more than
|
|
62
|
+
// DRIFT_US at the start of some sounds, and splitting there would cut a hole
|
|
63
|
+
// just after the sound begins.
|
|
64
|
+
const SILENCE_US = 50_000;
|
|
65
|
+
const DRIFT_US = 100_000;
|
|
66
|
+
const SETTLE_US = 500_000;
|
|
67
|
+
const framesOf = (packet) => Math.floor(packet.pcm.length / FRAME_BYTES);
|
|
68
|
+
const usOf = (frames) => (frames / SAMPLE_RATE) * 1e6;
|
|
69
|
+
// The packets in the order they arrived, which is the order they were played,
|
|
70
|
+
// cut into runs at each silence.
|
|
71
|
+
const runsOf = (packets) => {
|
|
72
|
+
const runs = [];
|
|
73
|
+
let run = [];
|
|
74
|
+
for (const packet of packets) {
|
|
75
|
+
const last = run.at(-1);
|
|
76
|
+
if (last &&
|
|
77
|
+
Math.abs(packet.timestampUs - (last.timestampUs + usOf(framesOf(last)))) >
|
|
78
|
+
SILENCE_US) {
|
|
79
|
+
runs.push(run);
|
|
80
|
+
run = [];
|
|
81
|
+
}
|
|
82
|
+
run.push(packet);
|
|
83
|
+
}
|
|
84
|
+
if (run.length)
|
|
85
|
+
runs.push(run);
|
|
86
|
+
return runs;
|
|
87
|
+
};
|
|
88
|
+
// Where a run's first frame belongs on a clock that starts at `startUs`, in
|
|
89
|
+
// microseconds: each packet's capture time less the audio before it in the
|
|
90
|
+
// run, and the median of those, which jitter and the first small packets do
|
|
91
|
+
// not move.
|
|
92
|
+
const anchorOf = (run, startUs) => {
|
|
93
|
+
let before = 0;
|
|
94
|
+
const offsets = run.map((packet) => {
|
|
95
|
+
const offset = packet.timestampUs - startUs - usOf(before);
|
|
96
|
+
before += framesOf(packet);
|
|
97
|
+
return offset;
|
|
98
|
+
});
|
|
99
|
+
offsets.sort((a, b) => a - b);
|
|
100
|
+
return offsets[Math.floor(offsets.length / 2)];
|
|
101
|
+
};
|
|
102
|
+
// The run split where a packet, past the run's first SETTLE_US of audio, has
|
|
103
|
+
// drifted more than DRIFT_US from the place the run gives it, so a long sound
|
|
104
|
+
// follows the clock: the part before the first such packet, and the rest.
|
|
105
|
+
const steady = (run, startUs) => {
|
|
106
|
+
const anchor = anchorOf(run, startUs);
|
|
107
|
+
let before = 0;
|
|
108
|
+
for (let i = 0; i < run.length; i++) {
|
|
109
|
+
const drift = run[i].timestampUs - startUs - (anchor + usOf(before));
|
|
110
|
+
if (usOf(before) > SETTLE_US && Math.abs(drift) > DRIFT_US)
|
|
111
|
+
return [run.slice(0, i), run.slice(i)];
|
|
112
|
+
before += framesOf(run[i]);
|
|
113
|
+
}
|
|
114
|
+
return [[...run], []];
|
|
115
|
+
};
|
|
52
116
|
// Lays audio packets on a timeline that starts at `startUs` and lasts
|
|
53
|
-
// `seconds`, as 16-bit stereo PCM
|
|
54
|
-
//
|
|
55
|
-
//
|
|
117
|
+
// `seconds`, as 16-bit stereo PCM, with silence where nothing arrived. Audio
|
|
118
|
+
// placed before the start is cut, and past the end is dropped; the result is
|
|
119
|
+
// exactly as long as asked. Where two runs overlap, the later one is heard.
|
|
56
120
|
export const assembleAudio = (packets, startUs, seconds) => {
|
|
57
121
|
const track = Buffer.alloc(Math.round(seconds * SAMPLE_RATE) * FRAME_BYTES);
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
122
|
+
const pending = runsOf(packets);
|
|
123
|
+
while (pending.length) {
|
|
124
|
+
const [run, rest] = steady(pending.shift(), startUs);
|
|
125
|
+
if (rest.length)
|
|
126
|
+
pending.unshift(rest);
|
|
127
|
+
let offset = Math.round((anchorOf(run, startUs) / 1e6) * SAMPLE_RATE);
|
|
128
|
+
for (const packet of run) {
|
|
129
|
+
const frames = framesOf(packet);
|
|
130
|
+
const from = Math.max(0, -offset);
|
|
131
|
+
const at = offset + from;
|
|
132
|
+
offset += frames;
|
|
133
|
+
if (from >= frames || at * FRAME_BYTES >= track.length)
|
|
134
|
+
continue;
|
|
135
|
+
packet.pcm.copy(track, at * FRAME_BYTES, from * FRAME_BYTES, frames * FRAME_BYTES);
|
|
65
136
|
}
|
|
66
|
-
if (from >= frames || offset * FRAME_BYTES >= track.length)
|
|
67
|
-
continue;
|
|
68
|
-
packet.pcm.copy(track, offset * FRAME_BYTES, from * FRAME_BYTES, frames * FRAME_BYTES);
|
|
69
137
|
}
|
|
70
138
|
return track;
|
|
71
139
|
};
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@fortemate/vega-vvd-driver",
|
|
3
|
-
"version": "0.2.
|
|
3
|
+
"version": "0.2.2",
|
|
4
4
|
"description": "Unofficial driver for the Vega Virtual Device: remote keys, screenshots and video with sound, from scripts and AI agents. CLI, Node library and MCP server.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"vega",
|