verikun 0.27.1 → 0.29.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/verikun/SKILL.md +11 -4
- package/CHANGELOG.md +39 -0
- package/dist/agent/remote.js +125 -18
- package/dist/cli.js +17 -2
- package/dist/device/failover.js +7 -2
- package/dist/drivers/adb.js +13 -7
- package/dist/install-timeouts.js +12 -0
- package/dist/server.js +229 -52
- package/dist/version.js +1 -1
- package/package.json +1 -1
|
@@ -514,6 +514,11 @@ installs again, warning on stderr that **that build's app data is gone**. A same
|
|
|
514
514
|
install still keeps its data. Do not treat the warning as a failure — it is how the
|
|
515
515
|
install succeeded. iOS has no equivalent recovery; the install just fails there.
|
|
516
516
|
|
|
517
|
+
**`vk install` also allows a version downgrade by default, when the build is debuggable.**
|
|
518
|
+
A build with a lower version code than what's already on the device installs normally
|
|
519
|
+
instead of failing with `INSTALL_FAILED_VERSION_DOWNGRADE`. A release-signed build is a
|
|
520
|
+
genuine exception — Android still refuses it, and that is not something to retry.
|
|
521
|
+
|
|
517
522
|
`vk server --devices all` (or `all-android` / `all-ios` / a serial list) serves a **pool**
|
|
518
523
|
from one address. Each run leases one device for its whole life, so a run's steps and
|
|
519
524
|
repairs always land on the same phone. `vk install --server` then installs on every device;
|
|
@@ -530,10 +535,12 @@ device that failed and is now on another one. What that means depends on the lin
|
|
|
530
535
|
the failure was real on A, and B has none of the state your flow built up. Start the
|
|
531
536
|
flow again from the top if you want it on B.
|
|
532
537
|
|
|
533
|
-
The server rules the bad device out
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
538
|
+
The server rules the bad device out; `vk devices --server <url>` shows why in its `NOTE`
|
|
539
|
+
column, and a pooled server re-adopts a device that comes back within a minute. A device
|
|
540
|
+
that is still attached keeps its place and is simply dealt last, so its own error keeps
|
|
541
|
+
reaching you rather than a bare "no device attached". A device that is **gone** leaves the
|
|
542
|
+
pool — so `capacity` can drop mid-job, and an install can come back `exit 0` having skipped
|
|
543
|
+
it. That is a success: nothing can be dealt a device running the previous build.
|
|
537
544
|
|
|
538
545
|
## The device is missing or wedged
|
|
539
546
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,45 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.29.0] - 2026-09-16
|
|
10
|
+
|
|
11
|
+
### Changed
|
|
12
|
+
- **`vk install`** (Android) allows a version downgrade by default (`adb install -d`),
|
|
13
|
+
instead of failing with `INSTALL_FAILED_VERSION_DOWNGRADE`, when the build is debuggable.
|
|
14
|
+
|
|
15
|
+
## [0.28.1] - 2026-09-15
|
|
16
|
+
|
|
17
|
+
### Fixed
|
|
18
|
+
- **`vk install --server`** sheds a pooled device whose install stalls, while the client now
|
|
19
|
+
honors its 15-minute request budget. ([#143])
|
|
20
|
+
|
|
21
|
+
### Changed
|
|
22
|
+
- **Docs rule**: a bug fix is changelog-only; `docs/` changes only when a contract moves or a page
|
|
23
|
+
went false. ([#140])
|
|
24
|
+
|
|
25
|
+
[#140]: https://github.com/ddikman/verikun/issues/140
|
|
26
|
+
[#143]: https://github.com/ddikman/verikun/issues/143
|
|
27
|
+
|
|
28
|
+
## [0.28.0] - 2026-09-14
|
|
29
|
+
|
|
30
|
+
A device that has gone now leaves a pooled `vk server` instead of being dealt out until
|
|
31
|
+
somebody notices.
|
|
32
|
+
|
|
33
|
+
### Fixed
|
|
34
|
+
- **`vk server --devices`** drops a device that is gone from the pool instead of dealing it
|
|
35
|
+
forever; the sweep readmits it when it returns. ([#139])
|
|
36
|
+
- **`vk install --server`** exits `0` when some pooled devices take the build, naming the rest
|
|
37
|
+
as `skipped`; only a build that fails everywhere is an error. ([#139])
|
|
38
|
+
- **A `--server` call that outlives Node's 300s fetch ceiling** now says so, and says the clock
|
|
39
|
+
was the client's, instead of a bare `fetch failed` the suite read as a dead device. ([#139])
|
|
40
|
+
|
|
41
|
+
### Changed
|
|
42
|
+
- **`/v1/health` `capacity` can now drop mid-job** on a pooled server, to `0`. A plain
|
|
43
|
+
`vk server` still keeps its only device and answers with that device's own error. ([#139])
|
|
44
|
+
- **`POST /v1/install`** answers `200 {devices, skipped}` where it used to answer `500` on a
|
|
45
|
+
partial failure. Older clients see a success with an unknown field. ([#139])
|
|
46
|
+
|
|
47
|
+
[#139]: https://github.com/ddikman/verikun/issues/139
|
|
9
48
|
## [0.27.1] - 2026-09-14
|
|
10
49
|
|
|
11
50
|
A hierarchy read the device killed is now waited out instead of ending the test.
|
package/dist/agent/remote.js
CHANGED
|
@@ -10,27 +10,98 @@
|
|
|
10
10
|
// local one. Recording stays a caller concern: this module never touches ./.verikun.
|
|
11
11
|
Object.defineProperty(exports, "__esModule", { value: true });
|
|
12
12
|
exports.describeStatus = describeStatus;
|
|
13
|
+
exports.transportReason = transportReason;
|
|
13
14
|
exports.pingServer = pingServer;
|
|
14
15
|
exports.remoteDeviceOp = remoteDeviceOp;
|
|
15
16
|
exports.remoteDeviceList = remoteDeviceList;
|
|
16
17
|
exports.createRemoteBackend = createRemoteBackend;
|
|
17
18
|
const node_fs_1 = require("node:fs");
|
|
18
19
|
const node_crypto_1 = require("node:crypto");
|
|
20
|
+
const node_http_1 = require("node:http");
|
|
21
|
+
const node_https_1 = require("node:https");
|
|
19
22
|
const node_path_1 = require("node:path");
|
|
20
23
|
const errors_1 = require("../errors");
|
|
24
|
+
const install_timeouts_1 = require("../install-timeouts");
|
|
25
|
+
const output_1 = require("../output");
|
|
21
26
|
const rpc_1 = require("../rpc");
|
|
27
|
+
/**
|
|
28
|
+
* The ceiling fetch-based calls below cannot exceed, whatever their own budgets say.
|
|
29
|
+
*
|
|
30
|
+
* Node's global `fetch` is undici, whose `headersTimeout` and `bodyTimeout` both default to
|
|
31
|
+
* 300s, and there is no dependency-free way to raise them: a `dispatcher` needs `undici`
|
|
32
|
+
* itself, which is bundled but not importable. Long-running install uploads therefore use
|
|
33
|
+
* Node's built-in http/https client instead; its explicit 15-minute timer is the real
|
|
34
|
+
* upload-and-response ceiling. The other AbortControllers remain a floor when set above 300s.
|
|
35
|
+
*
|
|
36
|
+
* MEASURED on Node v20.20.2 against a server that held its headers for 310s: the fetch
|
|
37
|
+
* rejected at 301s with `TypeError: fetch failed`, cause `HeadersTimeoutError`, code
|
|
38
|
+
* `UND_ERR_HEADERS_TIMEOUT`. The bare `fetch failed` is the whole problem — `describeStatus`
|
|
39
|
+
* never sees it, the suite reads the resulting exit 3 as the DEVICE being unreachable, and a
|
|
40
|
+
* healthy phone gets retired for a client-side clock. Named in `request` below so it says so.
|
|
41
|
+
*/
|
|
42
|
+
const FETCH_HEADERS_CEILING_MS = 300_000;
|
|
22
43
|
// Per-call ceilings. exec is generous: a single leaf may legitimately block for its
|
|
23
|
-
// whole auto-wait window or an explicit `wait --timeout`, plus device time.
|
|
44
|
+
// whole auto-wait window or an explicit `wait --timeout`, plus device time. Anything here
|
|
45
|
+
// above FETCH_HEADERS_CEILING_MS is aspirational unless it uses requestWithNodeHttp.
|
|
24
46
|
const HEALTH_TIMEOUT_MS = 10_000;
|
|
25
47
|
const ELEMENTS_TIMEOUT_MS = 60_000;
|
|
26
48
|
const EXEC_TIMEOUT_MS = 10 * 60_000;
|
|
27
|
-
const INSTALL_TIMEOUT_MS = 15 * 60_000;
|
|
28
49
|
const DEVICE_LIST_TIMEOUT_MS = 30_000;
|
|
29
|
-
//
|
|
30
|
-
// timed out rather than the client aborting
|
|
50
|
+
// Meant to sit above the server's own 4-minute boot ceiling, so the SERVER reports why a
|
|
51
|
+
// boot timed out rather than the client aborting first. It is exactly AT
|
|
52
|
+
// FETCH_HEADERS_CEILING_MS, so a boot that runs the full four minutes and then some is a
|
|
53
|
+
// photo finish — which is survivable only because `transportReason` now names the loser.
|
|
31
54
|
const DEVICE_START_TIMEOUT_MS = 5 * 60_000;
|
|
32
55
|
const DEVICE_STOP_TIMEOUT_MS = 60_000;
|
|
33
56
|
const trimUrl = (url) => url.replace(/\/+$/, '');
|
|
57
|
+
/**
|
|
58
|
+
* A dependency-free request path whose caller-owned timeout covers BOTH upload and response.
|
|
59
|
+
*
|
|
60
|
+
* Global fetch cannot wait past undici's fixed 300s header/body ceilings. An install may
|
|
61
|
+
* legitimately need 15 minutes, so it uses Node's native protocol clients instead of claiming
|
|
62
|
+
* a budget the transport cannot honor. Buffering the response preserves the Response boundary
|
|
63
|
+
* used below; install responses are small JSON descriptors, never artifacts.
|
|
64
|
+
*/
|
|
65
|
+
function requestWithNodeHttp(url, method, headers, body, timeoutMs) {
|
|
66
|
+
const endpoint = new URL(url);
|
|
67
|
+
const send = endpoint.protocol === 'http:' ? node_http_1.request : endpoint.protocol === 'https:' ? node_https_1.request : null;
|
|
68
|
+
if (!send)
|
|
69
|
+
return Promise.reject(new Error(`unsupported server protocol '${endpoint.protocol}'`));
|
|
70
|
+
const payload = body === undefined ? undefined : Buffer.isBuffer(body) ? body : Buffer.from(body);
|
|
71
|
+
const requestHeaders = {
|
|
72
|
+
...headers,
|
|
73
|
+
...(payload ? { 'content-length': String(payload.byteLength) } : {}),
|
|
74
|
+
};
|
|
75
|
+
return new Promise((resolve, reject) => {
|
|
76
|
+
let settled = false;
|
|
77
|
+
let timer;
|
|
78
|
+
const finish = (fn) => {
|
|
79
|
+
if (settled)
|
|
80
|
+
return;
|
|
81
|
+
settled = true;
|
|
82
|
+
if (timer)
|
|
83
|
+
clearTimeout(timer);
|
|
84
|
+
fn();
|
|
85
|
+
};
|
|
86
|
+
const req = send(endpoint, { method, headers: requestHeaders }, (res) => {
|
|
87
|
+
const chunks = [];
|
|
88
|
+
res.on('data', (chunk) => chunks.push(Buffer.from(chunk)));
|
|
89
|
+
res.on('error', (e) => finish(() => reject(e)));
|
|
90
|
+
res.on('aborted', () => finish(() => reject(new Error('the server aborted its response'))));
|
|
91
|
+
res.on('end', () => {
|
|
92
|
+
const status = res.statusCode ?? 500;
|
|
93
|
+
finish(() => resolve(new Response(Buffer.concat(chunks), { status })));
|
|
94
|
+
});
|
|
95
|
+
});
|
|
96
|
+
req.on('error', (e) => finish(() => reject(e)));
|
|
97
|
+
timer = setTimeout(() => {
|
|
98
|
+
const timeout = Object.assign(new Error('This operation was aborted'), { name: 'AbortError' });
|
|
99
|
+
req.destroy(timeout);
|
|
100
|
+
}, timeoutMs);
|
|
101
|
+
timer.unref?.();
|
|
102
|
+
req.end(payload);
|
|
103
|
+
});
|
|
104
|
+
}
|
|
34
105
|
/**
|
|
35
106
|
* Turn a non-2xx into the error the caller sees.
|
|
36
107
|
*
|
|
@@ -66,6 +137,28 @@ function describeStatus(status, body, url) {
|
|
|
66
137
|
}
|
|
67
138
|
return new errors_1.CliError(`verikun server error ${status} at ${url}${detail}`, exitCode);
|
|
68
139
|
}
|
|
140
|
+
/**
|
|
141
|
+
* Why the transport failed, in words an operator can act on. PURE — exported for the tests.
|
|
142
|
+
*
|
|
143
|
+
* The undici arm is the one that earns its keep. `fetch` reports its own header/body
|
|
144
|
+
* timeouts as a bare `TypeError: fetch failed` with the real cause one level down, and that
|
|
145
|
+
* string is indistinguishable from a server that is genuinely unreachable — which is how a
|
|
146
|
+
* five-minute install came to look like a dead phone, and how a lane came to be retired for
|
|
147
|
+
* it. Say which clock ran out, and say whose it was.
|
|
148
|
+
*/
|
|
149
|
+
function transportReason(e, timeoutMs) {
|
|
150
|
+
const ex = e;
|
|
151
|
+
if (ex?.name === 'AbortError')
|
|
152
|
+
return `timed out after ${Math.round(timeoutMs / 1000)}s`;
|
|
153
|
+
const code = ex?.cause?.code;
|
|
154
|
+
if (code === 'UND_ERR_HEADERS_TIMEOUT' || code === 'UND_ERR_BODY_TIMEOUT') {
|
|
155
|
+
const what = code === 'UND_ERR_HEADERS_TIMEOUT' ? 'send a response' : 'finish its response';
|
|
156
|
+
return (`the server did not ${what} within ${Math.round(FETCH_HEADERS_CEILING_MS / 1000)}s — ` +
|
|
157
|
+
"this is Node's own fetch ceiling on the CLIENT, not the device. " +
|
|
158
|
+
'The server may still be working; check its log before blaming the device');
|
|
159
|
+
}
|
|
160
|
+
return ex?.message ?? String(e);
|
|
161
|
+
}
|
|
69
162
|
async function readBody(res) {
|
|
70
163
|
try {
|
|
71
164
|
return (await res.json());
|
|
@@ -89,25 +182,31 @@ class RemoteTransport {
|
|
|
89
182
|
h.authorization = `Bearer ${this.opts.authKey}`;
|
|
90
183
|
return h;
|
|
91
184
|
}
|
|
92
|
-
async request(method, path, body, timeoutMs, extraHeaders = {}) {
|
|
185
|
+
async request(method, path, body, timeoutMs, extraHeaders = {}, transport = 'fetch') {
|
|
93
186
|
const url = `${this.base}${path}`;
|
|
94
|
-
|
|
95
|
-
const timer = setTimeout(() => controller.abort(), timeoutMs);
|
|
187
|
+
let timer;
|
|
96
188
|
let res;
|
|
97
189
|
try {
|
|
98
|
-
|
|
99
|
-
method,
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
190
|
+
if (transport === 'node-http') {
|
|
191
|
+
res = await requestWithNodeHttp(url, method, this.headers(extraHeaders), body, timeoutMs);
|
|
192
|
+
}
|
|
193
|
+
else {
|
|
194
|
+
const controller = new AbortController();
|
|
195
|
+
timer = setTimeout(() => controller.abort(), timeoutMs);
|
|
196
|
+
res = await fetch(url, {
|
|
197
|
+
method,
|
|
198
|
+
headers: this.headers(extraHeaders),
|
|
199
|
+
body,
|
|
200
|
+
signal: controller.signal,
|
|
201
|
+
});
|
|
202
|
+
}
|
|
104
203
|
}
|
|
105
204
|
catch (e) {
|
|
106
|
-
|
|
107
|
-
throw new errors_1.CliError(`cannot reach verikun server at ${url} (${reason})`, 3);
|
|
205
|
+
throw new errors_1.CliError(`cannot reach verikun server at ${url} (${transportReason(e, timeoutMs)})`, 3);
|
|
108
206
|
}
|
|
109
207
|
finally {
|
|
110
|
-
|
|
208
|
+
if (timer)
|
|
209
|
+
clearTimeout(timer);
|
|
111
210
|
}
|
|
112
211
|
if (!res.ok) {
|
|
113
212
|
const body = await readBody(res);
|
|
@@ -220,15 +319,23 @@ function createRemoteBackend(opts, health) {
|
|
|
220
319
|
throw new errors_1.CliError(`install: cannot read '${appPath}' (${e.message})`, 2);
|
|
221
320
|
}
|
|
222
321
|
const sha256 = (0, node_crypto_1.createHash)('sha256').update(buf).digest('hex');
|
|
223
|
-
const res = await t.request('POST', '/v1/install', buf,
|
|
322
|
+
const res = await t.request('POST', '/v1/install', buf, install_timeouts_1.REMOTE_INSTALL_TIMEOUT_MS, {
|
|
224
323
|
'content-type': 'application/octet-stream',
|
|
225
324
|
'x-verikun-ext': ext,
|
|
226
325
|
'x-verikun-sha256': sha256,
|
|
227
|
-
});
|
|
326
|
+
}, 'node-http');
|
|
228
327
|
// Install is the one operation the server replays elsewhere, so a move here means
|
|
229
328
|
// the build DID land — on a different device than the one we started with.
|
|
230
329
|
if (res.deviceChanged)
|
|
231
330
|
opts.onDeviceChange?.(res.deviceChanged);
|
|
331
|
+
// A PARTIAL install is a success, and it must not be a silent one: capacity just
|
|
332
|
+
// dropped, and the operator's next question is which phone to go and look at.
|
|
333
|
+
if (res.skipped?.length) {
|
|
334
|
+
(0, output_1.err)(`[verikun] server installed on ${(res.devices ?? []).join(', ') || '(none)'}; ` +
|
|
335
|
+
`${res.skipped.length} device(s) could not take this build and left the pool — ` +
|
|
336
|
+
res.skipped.map((s) => `${s.serial} (${s.reason})`).join('; '));
|
|
337
|
+
opts.onInstallSkipped?.(res.skipped);
|
|
338
|
+
}
|
|
232
339
|
},
|
|
233
340
|
async reset(appId) {
|
|
234
341
|
// Between-test housekeeping (vk suite): the step is deliberately NOT spliced
|
package/dist/cli.js
CHANGED
|
@@ -2269,10 +2269,12 @@ async function resolveBackend(platform, device, flags) {
|
|
|
2269
2269
|
// every recorded command, which beats a timer: it fires when work happens.
|
|
2270
2270
|
grant: (0, grant_1.processClaimGrant)(device, claims_1.releaseOwnClaims),
|
|
2271
2271
|
moves: [],
|
|
2272
|
+
skipped: [],
|
|
2272
2273
|
};
|
|
2273
2274
|
}
|
|
2274
2275
|
let runCtx = { platform, device };
|
|
2275
2276
|
const moves = [];
|
|
2277
|
+
const skipped = [];
|
|
2276
2278
|
/** Set by the last move; the preflight below reads it to decide whether re-asking is
|
|
2277
2279
|
* warranted, then clears it. */
|
|
2278
2280
|
let movedDuringCall;
|
|
@@ -2283,6 +2285,7 @@ async function resolveBackend(platform, device, flags) {
|
|
|
2283
2285
|
// is identical to a local run's. logStart travels from the server's device clock
|
|
2284
2286
|
// so archive-time / vk log scoping works without a local driver.
|
|
2285
2287
|
onStep: (step, artifacts, logStart) => run_1.Recorder.appendForeignStep(step, artifacts, { ...runCtx, logStart }),
|
|
2288
|
+
onInstallSkipped: (s) => skipped.push(...s),
|
|
2286
2289
|
onDeviceChange: (c) => {
|
|
2287
2290
|
moves.push(c);
|
|
2288
2291
|
movedDuringCall = c;
|
|
@@ -2361,6 +2364,7 @@ async function resolveBackend(platform, device, flags) {
|
|
|
2361
2364
|
grant: (0, grant_1.leaseGrant)(remote, serial),
|
|
2362
2365
|
remote: { url: server, version: health.version, reads },
|
|
2363
2366
|
moves,
|
|
2367
|
+
skipped,
|
|
2364
2368
|
};
|
|
2365
2369
|
}
|
|
2366
2370
|
/**
|
|
@@ -2659,7 +2663,7 @@ async function cmdInstall(positionals, flags) {
|
|
|
2659
2663
|
if (!(0, node_fs_1.existsSync)(path))
|
|
2660
2664
|
throw new errors_1.CliError(`install: '${appPath}' does not exist`, 2);
|
|
2661
2665
|
const platform = platformFromFlags(flags);
|
|
2662
|
-
const { backend, remote, moves } = await resolveBackend(platform, deviceFromFlags(flags, platform), flags);
|
|
2666
|
+
const { backend, remote, moves, skipped } = await resolveBackend(platform, deviceFromFlags(flags, platform), flags);
|
|
2663
2667
|
(0, output_1.err)(`[verikun] installing ${appPath}${remote ? ` via ${remote.url}` : ''}…`);
|
|
2664
2668
|
try {
|
|
2665
2669
|
await backend.install(path);
|
|
@@ -2676,11 +2680,22 @@ async function cmdInstall(positionals, flags) {
|
|
|
2676
2680
|
// device than the one the run started against, and a caller acting on the old serial
|
|
2677
2681
|
// (`adb -s … shell am start`) would be driving a phone without the build.
|
|
2678
2682
|
const moved = moves.length ? moves[moves.length - 1] : undefined;
|
|
2683
|
+
// A pooled server may have installed on some devices and dropped the rest. That is a
|
|
2684
|
+
// SUCCESS — the ones that missed the build are no longer leasable, so no later lane can
|
|
2685
|
+
// run the previous build and report green — but it is not a silent one: capacity changed.
|
|
2679
2686
|
if ((0, args_1.flagBool)(flags, 'json')) {
|
|
2680
|
-
(0, output_1.json)({
|
|
2687
|
+
(0, output_1.json)({
|
|
2688
|
+
installed: appPath,
|
|
2689
|
+
...(remote ? { server: remote.url } : {}),
|
|
2690
|
+
...(moved ? { deviceChanged: moved } : {}),
|
|
2691
|
+
...(skipped.length ? { skipped } : {}),
|
|
2692
|
+
});
|
|
2681
2693
|
}
|
|
2682
2694
|
else {
|
|
2683
2695
|
(0, output_1.out)(`installed ${appPath}${moved ? ` on ${moved.to}` : ''}`);
|
|
2696
|
+
if (skipped.length) {
|
|
2697
|
+
(0, output_1.out)(`skipped ${skipped.length} device(s), now out of the pool: ${skipped.map((s) => s.serial).join(', ')}`);
|
|
2698
|
+
}
|
|
2684
2699
|
}
|
|
2685
2700
|
return 0;
|
|
2686
2701
|
}
|
package/dist/device/failover.js
CHANGED
|
@@ -33,6 +33,7 @@ const ARTIFACT_RULES = [
|
|
|
33
33
|
const TRANSIENT_RULES = [[/INSTALL_FAILED_ABORTED/, 'the install session was aborted']];
|
|
34
34
|
// --- fast paths: name the reason and skip the probe. NEVER the gate. ----------
|
|
35
35
|
const UNREACHABLE_RULES = [
|
|
36
|
+
[/stopped responding.*install/i, 'the device stopped responding during install'],
|
|
36
37
|
[/device (?:'[^']*' )?not found/i, 'the device is not attached'],
|
|
37
38
|
[/no devices\/emulators found/i, 'no device is attached'],
|
|
38
39
|
[/device offline/i, 'the device is offline'],
|
|
@@ -49,6 +50,9 @@ const DEVICE_STATE_RULES = [
|
|
|
49
50
|
// `install` already tried removing it and reinstalling (drivers/adb.ts); reaching here
|
|
50
51
|
// means that did not work, so the wording must not imply nothing was attempted.
|
|
51
52
|
[/INSTALL_FAILED_UPDATE_INCOMPATIBLE/, 'a differently-signed build could not be replaced on the device'],
|
|
53
|
+
// `install` already retries with `-d` (drivers/adb.ts); reaching here means the build
|
|
54
|
+
// or device isn't debuggable, so Android is still refusing — a genuine per-device
|
|
55
|
+
// conflict, not a sign the retry never ran.
|
|
52
56
|
[/INSTALL_FAILED_VERSION_DOWNGRADE/, 'the device holds a newer build of this package'],
|
|
53
57
|
[/INSTALL_FAILED_ALREADY_EXISTS/, 'the package is already installed on the device'],
|
|
54
58
|
[/INSTALL_FAILED_DUPLICATE_PERMISSION/, 'another app on the device declares one of these permissions'],
|
|
@@ -166,8 +170,9 @@ function classifyInstallFailure(e) {
|
|
|
166
170
|
if (named)
|
|
167
171
|
return { move: true, kind: 'device-state', reason: named };
|
|
168
172
|
// The inversion. Not in the denylist ⇒ it is about the device, even though we have
|
|
169
|
-
// never seen this wording. Bounded by
|
|
170
|
-
// FIRST device's error on exhaustion, so a wrong guess costs time,
|
|
173
|
+
// never seen this wording. Bounded by the request-derived install hop limit, and the
|
|
174
|
+
// caller reports the FIRST device's error on exhaustion, so a wrong guess costs time,
|
|
175
|
+
// not diagnosis.
|
|
171
176
|
return { move: true, kind: 'device-state', reason: 'the device could not install this build', unclassified: true };
|
|
172
177
|
}
|
|
173
178
|
return classify(e, UNKNOWN_STAY);
|
package/dist/drivers/adb.js
CHANGED
|
@@ -1104,21 +1104,27 @@ class AdbDriver {
|
|
|
1104
1104
|
throw new errors_1.CliError(signatureConflictHelp(appPath, pkg, 'retry-failed', second), 3);
|
|
1105
1105
|
}
|
|
1106
1106
|
/**
|
|
1107
|
-
* One `adb install -r` attempt: null on success, else adb's collapsed output.
|
|
1107
|
+
* One `adb install -r -d` attempt: null on success, else adb's collapsed output.
|
|
1108
1108
|
*
|
|
1109
1109
|
* Returns rather than throws because the caller has to READ the failure to decide
|
|
1110
1110
|
* whether it is recoverable — and a thrown-and-caught CliError would put the final
|
|
1111
1111
|
* message's own prefix inside the message it is composing.
|
|
1112
1112
|
*
|
|
1113
1113
|
* `-r` reinstalls over an existing package keeping its data (the common
|
|
1114
|
-
* update-the-build-under-test case).
|
|
1115
|
-
*
|
|
1116
|
-
*
|
|
1117
|
-
*
|
|
1118
|
-
*
|
|
1114
|
+
* update-the-build-under-test case). `-d` allows a lower versionCode to replace a
|
|
1115
|
+
* higher one already on the device — unconditional here, not a retry like the
|
|
1116
|
+
* signature-conflict replace below, because it carries no data-loss trade-off (inert
|
|
1117
|
+
* when there is no downgrade, and keeps data the same way `-r` does when there is).
|
|
1118
|
+
* Android only honors it for a debuggable build, so a release-signed downgrade still
|
|
1119
|
+
* surfaces `INSTALL_FAILED_VERSION_DOWNGRADE` unresolved — see the comment on that
|
|
1120
|
+
* code in `device/failover.ts`. A large APK can legitimately take minutes to stream +
|
|
1121
|
+
* install, so the timeout is far above the 30s default. adb reports failures both as a
|
|
1122
|
+
* non-zero exit AND as a `Failure [REASON]` line on stdout with exit 0 (varies by adb
|
|
1123
|
+
* version) — check both, and require `Success` positively rather than merely inferring
|
|
1124
|
+
* it from the absence of `Failure`.
|
|
1119
1125
|
*/
|
|
1120
1126
|
tryInstall(appPath) {
|
|
1121
|
-
const r = (0, exec_1.runText)(ADB, this.withSerial(['install', '-r', appPath]), { timeout: 10 * 60 * 1000 });
|
|
1127
|
+
const r = (0, exec_1.runText)(ADB, this.withSerial(['install', '-r', '-d', appPath]), { timeout: 10 * 60 * 1000 });
|
|
1122
1128
|
const combined = `${r.stdout}\n${r.stderr}`;
|
|
1123
1129
|
if (r.code !== 0 || /^Failure\b/im.test(combined) || !/^Success\b/im.test(combined)) {
|
|
1124
1130
|
return combined.replace(/\s+/g, ' ').trim() || `exit code ${r.code}`;
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
"use strict";
|
|
2
|
+
// The remote-install time budget is shared by the client that owns the request and the
|
|
3
|
+
// server that may retry it. Keeping the numbers together prevents a server-side retry plan
|
|
4
|
+
// that can never finish before the client gives up.
|
|
5
|
+
Object.defineProperty(exports, "__esModule", { value: true });
|
|
6
|
+
exports.MAX_INSTALL_FAILOVER_HOPS = exports.INSTALL_DEVICE_TIMEOUT_MS = exports.REMOTE_INSTALL_TIMEOUT_MS = void 0;
|
|
7
|
+
/** Whole upload + install request, enforced by the remote client. */
|
|
8
|
+
exports.REMOTE_INSTALL_TIMEOUT_MS = 15 * 60_000;
|
|
9
|
+
/** One device gets four minutes; a measured ~170s large install still has useful margin. */
|
|
10
|
+
exports.INSTALL_DEVICE_TIMEOUT_MS = 4 * 60_000;
|
|
11
|
+
/** Three attempts consume at most 12m, leaving 3m of the request budget for the upload. */
|
|
12
|
+
exports.MAX_INSTALL_FAILOVER_HOPS = Math.max(0, Math.floor(exports.REMOTE_INSTALL_TIMEOUT_MS / exports.INSTALL_DEVICE_TIMEOUT_MS) - 1);
|
package/dist/server.js
CHANGED
|
@@ -45,6 +45,7 @@ const node_stream_1 = require("node:stream");
|
|
|
45
45
|
const promises_1 = require("node:stream/promises");
|
|
46
46
|
const args_1 = require("./args");
|
|
47
47
|
const errors_1 = require("./errors");
|
|
48
|
+
const install_timeouts_1 = require("./install-timeouts");
|
|
48
49
|
const drivers_1 = require("./drivers");
|
|
49
50
|
const manager_1 = require("./companion/manager");
|
|
50
51
|
const claims_1 = require("./device/claims");
|
|
@@ -86,10 +87,6 @@ const INSTALL_BODY_CAP = 512 * 1024 * 1024; // 512 MB app build
|
|
|
86
87
|
// to survive a client-side compile/repair pause, short enough that a crashed
|
|
87
88
|
// caller doesn't wedge the device.
|
|
88
89
|
const LOCK_IDLE_MS = 5 * 60 * 1000;
|
|
89
|
-
// How many times ONE request may move device. 2 moves = 3 devices tried, which sits
|
|
90
|
-
// comfortably inside the client's 15-minute install ceiling at ~1 minute an install,
|
|
91
|
-
// while a farm of ten wedged emulators cannot burn ten installs inside one request.
|
|
92
|
-
const MAX_FAILOVER_HOPS = 2;
|
|
93
90
|
// How often to ask whether the host's adb server has rotted. Generous on purpose: the
|
|
94
91
|
// check shells out to `log show` (~1s) and the condition it looks for accumulates over
|
|
95
92
|
// DAYS, so a tight interval would buy nothing and spend host time on every idle server.
|
|
@@ -108,6 +105,16 @@ const ADB_RECYCLE_CHECK_MS = 10 * 60 * 1000;
|
|
|
108
105
|
*/
|
|
109
106
|
const RECONCILE_INTERVAL_MS = 60_000;
|
|
110
107
|
const PROBE_RETRY_MS = 1000;
|
|
108
|
+
/** Kept distinct so an all-device timeout is never mistaken for evidence of a bad build. */
|
|
109
|
+
class InstallAttemptTimeoutError extends errors_1.CliError {
|
|
110
|
+
constructor(serial, timeoutMs) {
|
|
111
|
+
const duration = timeoutMs >= 1000 ? `${Math.round(timeoutMs / 1000)}s` : `${timeoutMs}ms`;
|
|
112
|
+
super(`device ${serial} stopped responding: install exceeded its ${duration} per-device deadline`, 3);
|
|
113
|
+
this.name = 'InstallAttemptTimeoutError';
|
|
114
|
+
}
|
|
115
|
+
}
|
|
116
|
+
/** Failover wraps an exhausted attempt in HttpError, so retain a wire-safe marker too. */
|
|
117
|
+
const isInstallAttemptTimeout = (e) => e instanceof InstallAttemptTimeoutError || (e instanceof Error && /per-device deadline/.test(e.message));
|
|
111
118
|
// Deliberately below the client's 5-minute ceiling, so a slow boot is reported by the
|
|
112
119
|
// side that knows WHY ("did not finish booting within 240s") rather than as a generic
|
|
113
120
|
// client-side abort.
|
|
@@ -519,7 +526,26 @@ function buildServer(config) {
|
|
|
519
526
|
: null;
|
|
520
527
|
// Like the reconcile timer: must never hold the process open at Ctrl-C.
|
|
521
528
|
recycleTimer?.unref?.();
|
|
522
|
-
|
|
529
|
+
/**
|
|
530
|
+
* Is this failure grounds to REMOVE the device from the pool, rather than merely deal it
|
|
531
|
+
* last? Two conditions, and both are necessary.
|
|
532
|
+
*
|
|
533
|
+
* `unreachable` is the only kind that qualifies, because it is the only one that says the
|
|
534
|
+
* device is not there. Every other kind describes a device that is present and unhappy —
|
|
535
|
+
* a full disk, a wedged app, an exit 3 nobody has classified — and for those, demotion
|
|
536
|
+
* plus recovery-by-traffic is right and this must not change: they can still produce the
|
|
537
|
+
* traffic that clears them. An absent device cannot, which is the whole defect (#139): the
|
|
538
|
+
* demotion is a sort key (`leaseFor`), so "dealt last" is still dealt, every round, and
|
|
539
|
+
* `restoreDevice` can never fire for a device that will never answer again.
|
|
540
|
+
*
|
|
541
|
+
* And only on a POOLED server, because only a pooled server sweeps. `reconcileOnce`
|
|
542
|
+
* returns immediately without `poolSpec` (`wantedSerials`) and its timer is never even
|
|
543
|
+
* created — see `ServerConfig.poolSpec`, "deliberately does not reconcile". Shedding
|
|
544
|
+
* where nothing readmits would trade a device that fails loudly for a server that is
|
|
545
|
+
* empty until someone restarts it: a worse failure, and a new one.
|
|
546
|
+
*/
|
|
547
|
+
const shedOnFailure = (kind) => kind === 'unreachable' && config.poolSpec !== undefined;
|
|
548
|
+
const pickFailoverDevice = (failed, reason, kind) => serializeFailover(() => pickFailoverDeviceLocked(failed, reason, kind));
|
|
523
549
|
/**
|
|
524
550
|
* Bring in a healthy replacement for `failed`. Returns the serial moved to, or null
|
|
525
551
|
* when none remains (which is not an error here — the caller reports the ORIGINAL
|
|
@@ -533,7 +559,7 @@ function buildServer(config) {
|
|
|
533
559
|
* only reports ready once its OWN `preflight()` has passed, so starting the worker IS
|
|
534
560
|
* the probe, run on the thread that will go on to use it.
|
|
535
561
|
*/
|
|
536
|
-
const pickFailoverDeviceLocked = async (failed, reason) => {
|
|
562
|
+
const pickFailoverDeviceLocked = async (failed, reason, kind) => {
|
|
537
563
|
const policy = config.failover;
|
|
538
564
|
if (!policy)
|
|
539
565
|
return null;
|
|
@@ -549,23 +575,29 @@ function buildServer(config) {
|
|
|
549
575
|
// see and where the server will actually go cannot drift. A pool member's own driver
|
|
550
576
|
// is not: it may be pointed at a corpse.
|
|
551
577
|
/**
|
|
552
|
-
* Nothing healthier exists.
|
|
553
|
-
*
|
|
578
|
+
* Nothing healthier exists. Decide what becomes of the failed device itself — THREE
|
|
579
|
+
* outcomes, not two, and which one applies is `shedOnFailure`'s question:
|
|
554
580
|
*
|
|
555
|
-
*
|
|
556
|
-
*
|
|
557
|
-
*
|
|
558
|
-
*
|
|
581
|
+
* - GONE, on a pooled server — shed it. It cannot serve and cannot recover by
|
|
582
|
+
* traffic, so leaving it in the pool means dealing it forever (#139). The sweep
|
|
583
|
+
* owns readmission, so capacity comes back on its own.
|
|
584
|
+
* - present but unhappy — demote it: worker, claim and slot kept, dealt last,
|
|
585
|
+
* restored by the first command that works.
|
|
586
|
+
* - already left on its own (its worker died) — nothing to remove, just clean up.
|
|
587
|
+
*
|
|
588
|
+
* The last two share a tail with the first, because "stop serving this device" has the
|
|
589
|
+
* same consequences however it came about.
|
|
559
590
|
*/
|
|
560
591
|
const shrink = async () => {
|
|
561
592
|
// A device whose worker DIED is already out of the pool, so there is nothing left to
|
|
562
593
|
// shed — but its holder still has to be evicted and its claim and companion handed
|
|
563
|
-
// back
|
|
564
|
-
//
|
|
594
|
+
// back. Asking whether it is still a member is what separates that case from a
|
|
595
|
+
// device we are removing ourselves.
|
|
565
596
|
const serving = pool.serials().includes(failed);
|
|
566
|
-
if (serving) {
|
|
567
|
-
// DEMOTE
|
|
568
|
-
// pool; it is simply dealt last until it does some work (see
|
|
597
|
+
if (serving && !shedOnFailure(kind)) {
|
|
598
|
+
// DEMOTE — the device is still THERE. It keeps its worker, its claim and its place
|
|
599
|
+
// in the pool; it is simply dealt last until it does some work (see
|
|
600
|
+
// `degradeDevice`). Contrast the shed below, which is only for a device that is not.
|
|
569
601
|
//
|
|
570
602
|
// This replaces "nothing healthier to move to — X left the pool". That rule read
|
|
571
603
|
// correctly on a SINGLE-device server, where it never actually fired (the last
|
|
@@ -573,10 +605,16 @@ function buildServer(config) {
|
|
|
573
605
|
// verdict — because a pool's own members are excluded from its candidate list, so
|
|
574
606
|
// "no candidate" is the normal case rather than the exceptional one. The argument
|
|
575
607
|
// for shedding was that continuing to hand out a broken device makes a pool a coin
|
|
576
|
-
// flip per lease; that is answered by ORDERING (a
|
|
577
|
-
// when nothing else is free), which costs no
|
|
578
|
-
//
|
|
579
|
-
//
|
|
608
|
+
// flip per lease; for a device that is PRESENT that is answered by ORDERING (a
|
|
609
|
+
// degraded device is chosen only when nothing else is free), which costs no
|
|
610
|
+
// capacity, and a caller that does reach it is better served by the truth about it
|
|
611
|
+
// than by a server that quietly halved.
|
|
612
|
+
//
|
|
613
|
+
// Ordering answers it only while the device can still come back, though. It cannot
|
|
614
|
+
// answer for a device that is GONE — "dealt last" is still dealt once the healthy
|
|
615
|
+
// devices are busy, which on a suite sized to the pool is every round, and no
|
|
616
|
+
// amount of ordering produces the traffic `restoreDevice` needs. That case is
|
|
617
|
+
// shed above, by `shedOnFailure`.
|
|
580
618
|
//
|
|
581
619
|
// The holder keeps its lease too: its device did not go anywhere, so there is no
|
|
582
620
|
// `deviceChanged` to send and nothing for the run to seal. The step that failed
|
|
@@ -585,7 +623,28 @@ function buildServer(config) {
|
|
|
585
623
|
degradeDevice(failed, reason);
|
|
586
624
|
return null;
|
|
587
625
|
}
|
|
588
|
-
|
|
626
|
+
if (serving) {
|
|
627
|
+
// SHED. The device is gone and this server sweeps, so removing it is not the
|
|
628
|
+
// one-way ratchet it was before the sweep existed (#114): `reconcileOnce` lists it
|
|
629
|
+
// as missing from what `--devices` asked for, retries with backoff, and
|
|
630
|
+
// `rejoinDevice` readmits it — bringing it up to `lastInstall` first — the moment
|
|
631
|
+
// it answers again. Capacity returns without anyone restarting anything.
|
|
632
|
+
//
|
|
633
|
+
// It keeps its QUARANTINE, unlike the demote branch above, and that asymmetry is
|
|
634
|
+
// the point: quarantine means "not serving, and ruled out", degradation means
|
|
635
|
+
// "serving but suspect", and the two are disjoint precisely so `/v1/health` and
|
|
636
|
+
// `exhaustedNote` can be read. A shed device genuinely is not serving, so it
|
|
637
|
+
// belongs in the same list as one whose worker died — which is the tail below,
|
|
638
|
+
// reached from here. `rejoinDevice` clears it on evidence, never on a clock.
|
|
639
|
+
//
|
|
640
|
+
// `degraded` must be given up though: it is defined as pool MEMBERS that recently
|
|
641
|
+
// failed, and a non-member left in it would have `/v1/health` reporting a device it
|
|
642
|
+
// no longer serves, in a list whose whole meaning is that it still does.
|
|
643
|
+
pool.retire(failed);
|
|
644
|
+
degraded.delete(failed);
|
|
645
|
+
(0, output_1.err)(`[server] pool: ${failed} left the pool — ${reason} (the sweep readmits it when it answers again)`);
|
|
646
|
+
}
|
|
647
|
+
// NOT serving — its worker died, or the shed above just removed it, so there is
|
|
589
648
|
// nothing to demote. The holder is EVICTED, not migrated: without a replacement there
|
|
590
649
|
// is no `deviceChanged` to send, so the client never learns to seal its run — and
|
|
591
650
|
// merely dropping the lease would let its next request silently draw some other device
|
|
@@ -690,7 +749,14 @@ function buildServer(config) {
|
|
|
690
749
|
* Is this device actually gone? Two probes a second apart, because that gap is the only
|
|
691
750
|
* thing separating a USB re-enumeration or a mid-`launch --clear` gap from a dead box —
|
|
692
751
|
* and quarantining a healthy device is the expensive mistake here. Returns the reason
|
|
693
|
-
* when dead, undefined when it was a blip.
|
|
752
|
+
* AND the probe's own verdict kind when dead, undefined when it was a blip.
|
|
753
|
+
*
|
|
754
|
+
* The kind is carried out because the probe is often the better-classified of the two
|
|
755
|
+
* failures. The operation that brought us here may have failed with a string nothing
|
|
756
|
+
* recognises (an unclassified exit 3, which is what earns a probe in the first place),
|
|
757
|
+
* while `preflight` on a detached phone says `device '<serial>' not found` — the exact
|
|
758
|
+
* `UNREACHABLE_RULES` wording. Reporting the ORIGINAL verdict's kind there would decide
|
|
759
|
+
* "shed or demote" from the vaguer of two answers about the same device.
|
|
694
760
|
*/
|
|
695
761
|
const deviceIsDead = async (handle) => {
|
|
696
762
|
let last = '';
|
|
@@ -722,7 +788,7 @@ function buildServer(config) {
|
|
|
722
788
|
(0, output_1.err)(`[server] probe on ${handle.serial}: ${verdict.reason} (${verdict.kind}) — a host problem, not this device`);
|
|
723
789
|
return undefined;
|
|
724
790
|
}
|
|
725
|
-
return last || 'the device stopped answering';
|
|
791
|
+
return { reason: last || 'the device stopped answering', kind: verdict.kind };
|
|
726
792
|
};
|
|
727
793
|
/**
|
|
728
794
|
* A non-install operation failed. Move off the device if it is genuinely at fault —
|
|
@@ -761,6 +827,7 @@ function buildServer(config) {
|
|
|
761
827
|
try {
|
|
762
828
|
const verdict = (0, failover_1.classifyFailure)(e);
|
|
763
829
|
let reason = verdict.reason;
|
|
830
|
+
let kind = verdict.kind;
|
|
764
831
|
if (!verdict.move) {
|
|
765
832
|
// Only an unrecognised exit 3 earns a probe; `transient` and `toolchain` set
|
|
766
833
|
// probe:false precisely so a mid-launch gap or a missing adb cannot become a move.
|
|
@@ -777,15 +844,21 @@ function buildServer(config) {
|
|
|
777
844
|
(0, output_1.err)(`[server] ${what}: ${from} failed but probes healthy — staying (${verdict.reason})`);
|
|
778
845
|
return undefined; // a blip — the test rerun is the right answer, not a new device
|
|
779
846
|
}
|
|
780
|
-
reason = dead;
|
|
847
|
+
reason = dead.reason;
|
|
848
|
+
// The PROBE's verdict, not the original failure's. We are here because the
|
|
849
|
+
// operation failed with something nothing recognised; `preflight` on a detached
|
|
850
|
+
// phone says `device '<serial>' not found`, which is classified. Taking the vaguer
|
|
851
|
+
// of two answers about the same device is how a detachment that first showed up as
|
|
852
|
+
// an odd exit 3 would be demoted forever instead of shed.
|
|
853
|
+
kind = dead.kind;
|
|
781
854
|
noteVerdict({ ...verdict, move: true }, e, what);
|
|
782
855
|
}
|
|
783
856
|
(0, output_1.err)(`[server] ${what}: FAILED on ${from} — ${reason}`);
|
|
784
857
|
quarantineDevice(from, reason);
|
|
785
|
-
// pickFailoverDevice has already said which
|
|
786
|
-
//
|
|
787
|
-
//
|
|
788
|
-
const to = await pickFailoverDevice(from, reason);
|
|
858
|
+
// pickFailoverDevice has already said which no-move outcome happened — the device
|
|
859
|
+
// was shed, or demoted, or had already left. A second line here would contradict
|
|
860
|
+
// one of them.
|
|
861
|
+
const to = await pickFailoverDevice(from, reason, kind);
|
|
789
862
|
if (!to)
|
|
790
863
|
return undefined;
|
|
791
864
|
return { from, to, reason, retried: false };
|
|
@@ -958,6 +1031,9 @@ function buildServer(config) {
|
|
|
958
1031
|
return new server_http_1.HttpError(409, 'the device this run was using left the pool and nothing healthy replaced it — ' +
|
|
959
1032
|
'start a fresh run; this one cannot continue on another device', 3);
|
|
960
1033
|
}
|
|
1034
|
+
// An empty pool never reaches here: the deviceless guard in the router answers 503 for
|
|
1035
|
+
// every route that takes a lease, and it names `lostDevice` while doing it. So `n` is
|
|
1036
|
+
// always >= 1 and this only ever describes CONTENTION, which is what 409 means.
|
|
961
1037
|
const n = pool.serials().length;
|
|
962
1038
|
return new server_http_1.HttpError(409, n > 1
|
|
963
1039
|
? `all ${n} devices are leased by other active runs — retry when one finishes`
|
|
@@ -1182,7 +1258,7 @@ function buildServer(config) {
|
|
|
1182
1258
|
* On exhaustion it throws the FIRST device's error, never the last. That inversion is
|
|
1183
1259
|
* what makes move-by-default safe: a wrong move costs time, not the diagnosis.
|
|
1184
1260
|
*/
|
|
1185
|
-
async function installWithFailover(serial, tmpPath) {
|
|
1261
|
+
async function installWithFailover(serial, tmpPath, timedOut) {
|
|
1186
1262
|
let change;
|
|
1187
1263
|
let moves = 0;
|
|
1188
1264
|
let firstError;
|
|
@@ -1197,9 +1273,9 @@ function buildServer(config) {
|
|
|
1197
1273
|
* wrapper keeps no result on a throw), so dropping it leaves the operator holding a
|
|
1198
1274
|
* serial the server has already left.
|
|
1199
1275
|
*/
|
|
1200
|
-
const hopOrThrow = async (why, giveUp) => {
|
|
1276
|
+
const hopOrThrow = async (why, giveUp, kind) => {
|
|
1201
1277
|
quarantineDevice(from, why);
|
|
1202
|
-
const to = await pickFailoverDevice(from, why);
|
|
1278
|
+
const to = await pickFailoverDevice(from, why, kind);
|
|
1203
1279
|
if (to === null) {
|
|
1204
1280
|
// A pool that emptied with nothing having moved keeps its own 503 — a more
|
|
1205
1281
|
// accurate status than a wrapped 500.
|
|
@@ -1222,13 +1298,36 @@ function buildServer(config) {
|
|
|
1222
1298
|
// are still on our disk), so it hops rather than throwing a bare 503 that never
|
|
1223
1299
|
// reaches the failover machinery at all.
|
|
1224
1300
|
const gone = firstError ?? new server_http_1.HttpError(503, `device ${from} is no longer attached`, 3);
|
|
1225
|
-
if (!config.failover || hop >=
|
|
1301
|
+
if (!config.failover || hop >= install_timeouts_1.MAX_INSTALL_FAILOVER_HOPS)
|
|
1226
1302
|
throw gone;
|
|
1227
|
-
|
|
1303
|
+
// `unreachable` is the literal truth — it is not in the pool — and it is also
|
|
1304
|
+
// inert here: `shrink` sees a non-member and takes its cleanup tail either way.
|
|
1305
|
+
from = await hopOrThrow('the device left the pool mid-install', gone, 'unreachable');
|
|
1228
1306
|
continue;
|
|
1229
1307
|
}
|
|
1230
1308
|
try {
|
|
1231
|
-
|
|
1309
|
+
const timeoutMs = config.installAttemptMs ?? install_timeouts_1.INSTALL_DEVICE_TIMEOUT_MS;
|
|
1310
|
+
await new Promise((resolve, reject) => {
|
|
1311
|
+
let settled = false;
|
|
1312
|
+
const finish = (fn) => {
|
|
1313
|
+
if (settled)
|
|
1314
|
+
return;
|
|
1315
|
+
settled = true;
|
|
1316
|
+
clearTimeout(timer);
|
|
1317
|
+
fn();
|
|
1318
|
+
};
|
|
1319
|
+
const timer = setTimeout(() => {
|
|
1320
|
+
const timeout = new InstallAttemptTimeoutError(from, timeoutMs);
|
|
1321
|
+
// A race alone would leave the worker blocked in adb forever. Retiring it is
|
|
1322
|
+
// what releases the transport and makes the deadline a recovery mechanism.
|
|
1323
|
+
timedOut.add(from);
|
|
1324
|
+
quarantineDevice(from, timeout.message);
|
|
1325
|
+
pool.retire(from);
|
|
1326
|
+
finish(() => reject(timeout));
|
|
1327
|
+
}, timeoutMs);
|
|
1328
|
+
timer.unref?.();
|
|
1329
|
+
void handle.install(tmpPath).then(() => finish(resolve), (e) => finish(() => reject(e)));
|
|
1330
|
+
});
|
|
1232
1331
|
return { change, moves };
|
|
1233
1332
|
}
|
|
1234
1333
|
catch (e) {
|
|
@@ -1239,9 +1338,19 @@ function buildServer(config) {
|
|
|
1239
1338
|
noteVerdict(verdict, e, 'install');
|
|
1240
1339
|
// The artifact is broken / the caller is wrong / failover is off / we are out of
|
|
1241
1340
|
// hops: report the first failure unchanged, exactly as before this feature.
|
|
1242
|
-
if (!verdict.move || !config.failover || hop >=
|
|
1341
|
+
if (!verdict.move || !config.failover || hop >= install_timeouts_1.MAX_INSTALL_FAILOVER_HOPS) {
|
|
1342
|
+
// A timeout retired its worker before this catch. If it cannot move, finish the
|
|
1343
|
+
// cleanup that pickFailoverDevice would otherwise own; especially, never leave
|
|
1344
|
+
// a host-global claim pinned to a worker that no longer exists.
|
|
1345
|
+
if (isInstallAttemptTimeout(e)) {
|
|
1346
|
+
evictHoldersOf(from, `${from} left the pool after its install timed out`);
|
|
1347
|
+
(0, manager_1.releaseCompanionOn)(from);
|
|
1348
|
+
if ((0, claims_1.claimsEnabled)(claimEnv))
|
|
1349
|
+
(0, claims_1.releaseClaim)(from, { ...claimOpts, mineOnly: true });
|
|
1350
|
+
}
|
|
1243
1351
|
throw firstError;
|
|
1244
|
-
|
|
1352
|
+
}
|
|
1353
|
+
from = await hopOrThrow(verdict.reason, firstError, verdict.kind);
|
|
1245
1354
|
}
|
|
1246
1355
|
}
|
|
1247
1356
|
}
|
|
@@ -1290,9 +1399,10 @@ function buildServer(config) {
|
|
|
1290
1399
|
throw new server_http_1.HttpError(503, `no device is left to install onto${lostDevice ? ` — last loss: ${lostDevice}` : ''}`, 3);
|
|
1291
1400
|
}
|
|
1292
1401
|
(0, output_1.err)(`[server] install: received ${size} bytes (.${ext}), installing on ${targets.join(', ')}…`);
|
|
1402
|
+
const timedOut = new Set();
|
|
1293
1403
|
const outcomes = await Promise.all(targets.map(async (serial) => {
|
|
1294
1404
|
try {
|
|
1295
|
-
return { serial, ...(await installWithFailover(serial, tmpPath)), error: null };
|
|
1405
|
+
return { serial, ...(await installWithFailover(serial, tmpPath, timedOut)), error: null };
|
|
1296
1406
|
}
|
|
1297
1407
|
catch (e) {
|
|
1298
1408
|
return { serial, change: undefined, moves: 0, error: e };
|
|
@@ -1300,6 +1410,12 @@ function buildServer(config) {
|
|
|
1300
1410
|
}));
|
|
1301
1411
|
const failed = outcomes.filter((o) => o.error);
|
|
1302
1412
|
const moved = outcomes.filter((o) => o.change);
|
|
1413
|
+
// Where each outcome ENDED UP. An install that moved ran on its replacement, not on
|
|
1414
|
+
// the serial it started from, so the device that holds this build — or conspicuously
|
|
1415
|
+
// does not — is the last one it was on, never `o.serial`.
|
|
1416
|
+
const landedOn = (o) => o.change?.to ?? o.serial;
|
|
1417
|
+
const installed = outcomes.filter((o) => !o.error).map(landedOn);
|
|
1418
|
+
let skipped = [];
|
|
1303
1419
|
if (failed.length) {
|
|
1304
1420
|
// One artifact, many devices: if it failed everywhere the file is the suspect, so
|
|
1305
1421
|
// surface the FIRST device's error unchanged rather than a summary that buries it.
|
|
@@ -1310,32 +1426,91 @@ function buildServer(config) {
|
|
|
1310
1426
|
// polarity, since the device side is open-ended and OEM-specific. That is right
|
|
1311
1427
|
// for ONE device failing; applied to every device at once it condemns the whole
|
|
1312
1428
|
// pool for what this very branch has just concluded is a bad build. Undo them.
|
|
1313
|
-
const
|
|
1314
|
-
if (
|
|
1315
|
-
|
|
1429
|
+
const hasTimeout = timedOut.size > 0;
|
|
1430
|
+
if (!hasTimeout) {
|
|
1431
|
+
// No lane succeeded and none hung: the common input is now the stronger
|
|
1432
|
+
// suspect, so undo the per-device quarantines made by the move-by-default
|
|
1433
|
+
// classifier.
|
|
1434
|
+
const condemned = targets.filter((t) => quarantine.delete(t));
|
|
1435
|
+
if (condemned.length) {
|
|
1436
|
+
(0, output_1.err)(`[server] install: failed on every device, so the build is the suspect — un-quarantining ${condemned.join(', ')}`);
|
|
1437
|
+
}
|
|
1438
|
+
}
|
|
1439
|
+
else {
|
|
1440
|
+
// A deadline is direct evidence about a device, not the artifact. Preserve
|
|
1441
|
+
// every quarantine when any lane timed out; otherwise the next request could
|
|
1442
|
+
// immediately be dealt the same wedged worker again.
|
|
1443
|
+
(0, output_1.err)('[server] install: every device failed and at least one timed out — keeping the device quarantines');
|
|
1316
1444
|
}
|
|
1317
1445
|
throw failed[0].error;
|
|
1318
1446
|
}
|
|
1319
|
-
//
|
|
1320
|
-
//
|
|
1321
|
-
//
|
|
1322
|
-
|
|
1447
|
+
// PARTIAL. Two healthy phones took the build and one did not. Answering 500 for the
|
|
1448
|
+
// whole pool is what turned one detached device into a dead CI job (#139) — and it
|
|
1449
|
+
// dies at the install step, so the run has already paid for an app build and tested
|
|
1450
|
+
// nothing.
|
|
1451
|
+
//
|
|
1452
|
+
// What made the 500 defensible is the fan-out's own rule, one line up: a lane dealt
|
|
1453
|
+
// a device that missed this build runs the PREVIOUS one and reports green, which is
|
|
1454
|
+
// the worst result this server can produce. The answer is not to soften that rule
|
|
1455
|
+
// but to SATISFY it — a device that did not take the build leaves the pool, so no
|
|
1456
|
+
// lease can reach it. `rejoinDevice` already makes exactly this call out loud
|
|
1457
|
+
// ("serving the wrong build is worse than not serving") and offers the same remedy:
|
|
1458
|
+
// the sweep readmits it and installs `lastInstall` before it is dealt any work.
|
|
1459
|
+
//
|
|
1460
|
+
// Done regardless of `config.failover`. The kill switch governs MOVING BETWEEN
|
|
1461
|
+
// devices; it was never a licence to serve a stale build, and the sweep that brings
|
|
1462
|
+
// the device back is gated on `poolSpec`, not on failover.
|
|
1463
|
+
skipped = failed.map((f) => {
|
|
1464
|
+
const serial = landedOn(f);
|
|
1465
|
+
const reason = (0, server_http_1.firstLine)(f.error.message);
|
|
1466
|
+
const serving = pool.serials().includes(serial);
|
|
1467
|
+
if (serving)
|
|
1468
|
+
pool.retire(serial);
|
|
1469
|
+
// Same two rules as the shed in `shrink`: `degraded` is for MEMBERS, and a
|
|
1470
|
+
// device that is not serving belongs in `quarantine` — which `rejoinDevice`
|
|
1471
|
+
// clears on the evidence of a worker that started and a build that installed.
|
|
1472
|
+
// The timeout path retired its worker already, but still needs this common
|
|
1473
|
+
// bookkeeping tail. Other already-absent failures keep their existing handling.
|
|
1474
|
+
if (serving || isInstallAttemptTimeout(f.error)) {
|
|
1475
|
+
degraded.delete(serial);
|
|
1476
|
+
quarantineDevice(serial, `did not take the current build — ${reason}`);
|
|
1477
|
+
evictHoldersOf(serial, `${serial} left the pool without the current build`);
|
|
1478
|
+
(0, manager_1.releaseCompanionOn)(serial);
|
|
1479
|
+
if ((0, claims_1.claimsEnabled)(claimEnv))
|
|
1480
|
+
(0, claims_1.releaseClaim)(serial, { ...claimOpts, mineOnly: true });
|
|
1481
|
+
}
|
|
1482
|
+
return { serial, reason };
|
|
1483
|
+
});
|
|
1484
|
+
(0, output_1.err)(`[server] install: partial — ${installed.join(', ')} took the build; ` +
|
|
1485
|
+
`removed from the pool: ${skipped.map((s) => `${s.serial} (${s.reason})`).join('; ')}`);
|
|
1323
1486
|
}
|
|
1324
1487
|
for (const m of moved)
|
|
1325
1488
|
(0, output_1.err)(`[server] install: ${m.serial} → ${m.change.to} (${m.moves} move(s))`);
|
|
1326
|
-
(0, output_1.err)(`[server] install: done on ${
|
|
1489
|
+
(0, output_1.err)(`[server] install: done on ${installed.join(', ')}`);
|
|
1327
1490
|
// Retain the artifact so a device that rejoins later can be brought up to this build
|
|
1328
1491
|
// (see `rejoinDevice`). Renamed out of the per-request temp name into one stable slot,
|
|
1329
1492
|
// so at most one build is ever held and each install replaces the last.
|
|
1493
|
+
//
|
|
1494
|
+
// Reached on a PARTIAL install too, and load-bearing there: the devices just removed
|
|
1495
|
+
// are precisely the ones the sweep will readmit, and `rejoinDevice` brings a returning
|
|
1496
|
+
// device up to `lastInstall`. Retaining only on a clean sweep would hand each of them
|
|
1497
|
+
// the PREVIOUS build on the way back in — and `rejoinDevice`'s own check would pass,
|
|
1498
|
+
// because an install that succeeds is all it can see.
|
|
1330
1499
|
retainInstall(tmpPath, ext);
|
|
1331
1500
|
retained = true;
|
|
1501
|
+
// Only a move whose destination SURVIVED is worth reporting. The client re-points its
|
|
1502
|
+
// run context on `deviceChanged`, so naming a device the partial branch retired three
|
|
1503
|
+
// lines ago would send its next step to a serial this server no longer serves.
|
|
1504
|
+
const survivors = new Set(pool.serials());
|
|
1505
|
+
const reportableMove = moved.map((m) => m.change).find((c) => survivors.has(c.to));
|
|
1332
1506
|
const body = {
|
|
1333
1507
|
ok: true,
|
|
1334
1508
|
bytes: size,
|
|
1335
1509
|
sha256: digest,
|
|
1336
|
-
devices:
|
|
1510
|
+
devices: installed,
|
|
1511
|
+
...(skipped.length ? { skipped } : {}),
|
|
1337
1512
|
// The wire field is singular; a pool that moved more than one device logs the rest.
|
|
1338
|
-
...(
|
|
1513
|
+
...(reportableMove ? { deviceChanged: reportableMove } : {}),
|
|
1339
1514
|
};
|
|
1340
1515
|
(0, server_http_1.sendJson)(res, 200, body);
|
|
1341
1516
|
}
|
|
@@ -1360,10 +1535,12 @@ function buildServer(config) {
|
|
|
1360
1535
|
* phone another job is mid-test on. Refusing plainly beats a rule nobody can predict.
|
|
1361
1536
|
* The GET listing stays available, because reading what is attached is safe.
|
|
1362
1537
|
*
|
|
1363
|
-
*
|
|
1364
|
-
*
|
|
1365
|
-
*
|
|
1366
|
-
*
|
|
1538
|
+
* This used to carry a known cost — that it was the ONLY thing clearing a quarantine, so
|
|
1539
|
+
* on a pool a device ruled out by failover stayed out until the server was restarted.
|
|
1540
|
+
* That is no longer true and the refusal no longer says it: `rejoinDevice` clears the
|
|
1541
|
+
* quarantine (and `degraded`, and `failedOver`) when the sweep readmits a device, on the
|
|
1542
|
+
* evidence of a worker that started and a build that installed. The refusal itself stands
|
|
1543
|
+
* — power-cycling one member of a pool another job is mid-test on is what it prevents.
|
|
1367
1544
|
*/
|
|
1368
1545
|
function requireSingleDevice(op) {
|
|
1369
1546
|
const n = pool.serials().length;
|
package/dist/version.js
CHANGED
|
@@ -3,4 +3,4 @@ Object.defineProperty(exports, "__esModule", { value: true });
|
|
|
3
3
|
exports.VERSION = void 0;
|
|
4
4
|
// GENERATED by scripts/gen-version.mjs from package.json's "version" at build time
|
|
5
5
|
// (the `prebuild` script). Do NOT edit by hand; bump package.json instead.
|
|
6
|
-
exports.VERSION = '0.
|
|
6
|
+
exports.VERSION = '0.29.0';
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "verikun",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.29.0",
|
|
4
4
|
"description": "Drive Android emulators/devices and iOS simulators for AI agents: tap, type, swipe, screenshot, and inspect the UI hierarchy by semantic identifiers — like Puppeteer for native apps.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"android",
|