verikun 0.28.0 → 0.29.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -514,6 +514,11 @@ installs again, warning on stderr that **that build's app data is gone**. A same
514
514
  install still keeps its data. Do not treat the warning as a failure — it is how the
515
515
  install succeeded. iOS has no equivalent recovery; the install just fails there.
516
516
 
517
+ **`vk install` also allows a version downgrade by default, when the build is debuggable.**
518
+ A build with a lower version code than what's already on the device installs normally
519
+ instead of failing with `INSTALL_FAILED_VERSION_DOWNGRADE`. A release-signed build is a
520
+ genuine exception — Android still refuses it, and that is not something to retry.
521
+
517
522
  `vk server --devices all` (or `all-android` / `all-ios` / a serial list) serves a **pool**
518
523
  from one address. Each run leases one device for its whole life, so a run's steps and
519
524
  repairs always land on the same phone. `vk install --server` then installs on every device;
package/CHANGELOG.md CHANGED
@@ -6,6 +6,33 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.29.1] - 2026-09-26
10
+
11
+ ### Fixed
12
+ - **`vk ai`** no longer guesses `id:` selectors; an element the test names only by label compiles
13
+ to `text:` — write the id into the test to pin it. ([#148])
14
+
15
+ [#148]: https://github.com/ddikman/verikun/issues/148
16
+
17
+ ## [0.29.0] - 2026-09-16
18
+
19
+ ### Changed
20
+ - **`vk install`** (Android) allows a version downgrade by default (`adb install -d`),
21
+ instead of failing with `INSTALL_FAILED_VERSION_DOWNGRADE`, when the build is debuggable.
22
+
23
+ ## [0.28.1] - 2026-09-15
24
+
25
+ ### Fixed
26
+ - **`vk install --server`** sheds a pooled device whose install stalls, while the client now
27
+ honors its 15-minute request budget. ([#143])
28
+
29
+ ### Changed
30
+ - **Docs rule**: a bug fix is changelog-only; `docs/` changes only when a contract moves or a page
31
+ went false. ([#140])
32
+
33
+ [#140]: https://github.com/ddikman/verikun/issues/140
34
+ [#143]: https://github.com/ddikman/verikun/issues/143
35
+
9
36
  ## [0.28.0] - 2026-09-14
10
37
 
11
38
  A device that has gone now leaves a pooled `vk server` instead of being dealt out until
@@ -26,7 +53,6 @@ somebody notices.
26
53
  partial failure. Older clients see a success with an unknown field. ([#139])
27
54
 
28
55
  [#139]: https://github.com/ddikman/verikun/issues/139
29
-
30
56
  ## [0.27.1] - 2026-09-14
31
57
 
32
58
  A hierarchy read the device killed is now waited out instead of ending the test.
@@ -118,7 +118,7 @@ Do NOT hard-code a run of indices you were not told the length of. If the prose
118
118
  varies per run — emit a while-present over {{ctx.i}} rather than tap _0, _1, _2, _3.
119
119
  A hard-coded list is right only when the prose states the exact count.
120
120
 
121
- SELECTORS (the engine auto-heals case/whitespace/partial, so prefer stable identifiers):
121
+ SELECTORS (the engine auto-heals case/whitespace/partial, so prefer an id the test gives you):
122
122
  @login resource-id 'login' (shorthand for id:login)
123
123
  id:login resource-id (full, suffix, or short)
124
124
  text:Sign in visible text (case-insensitive)
@@ -156,7 +156,17 @@ RULES:
156
156
  below the fold — that is now redundant. Emit an explicit swipe only when the SCROLLING
157
157
  ITSELF is what the test asks for ("scroll the feed three times"), or to reveal content
158
158
  that is not in the hierarchy until it is built (an infinite/lazy list).
159
- - Prefer resource-id / accessibility selectors over visible text where possible.
159
+ - IDS COME FROM THE TEST. Use @x / id:x only with an id the test writes (or one a PRIOR PLAN
160
+ shown to you already uses). You cannot see the app, so never compose an id from a label, a
161
+ description, or the pattern of the app's other ids: a made-up id matches nothing, and a
162
+ wait/assert on it FAILS the test with no repair. An element the test names only by what it
163
+ says ("tap Sign in", "the Home tab") is text:<those words> — forgiving, and it also matches an
164
+ accessibility label.
165
+ - An outcome the test states ("the tab bar (Home, Search, Profile) appears") is checked on what
166
+ it NAMES — the id it gives, or the words it quotes or lists — never on an id you made up. A
167
+ description that names nothing you could select without inventing it ("it shows the
168
+ signed-in user's name") is context, not a step; so is rationale (why a step exists, what a
169
+ field already holds).
160
170
  - Translate the test literally and minimally: do not invent ACTION steps (tap/text/swipe/key/assert)
161
171
  the prose does not imply. The ONE exception is screenshot — insert screenshot steps liberally as
162
172
  post-run review evidence: after each screen transition (launch, a navigation tap, a submit, a
@@ -2,7 +2,7 @@
2
2
  // Compile-fidelity lint: does the plan the model produced still say what the prose said?
3
3
  //
4
4
  // The model is a compiler, and this is the compiler's own sanity check. It exists because
5
- // compilation is NONDETERMINISTIC in a way that is invisible until a run fails, and three
5
+ // compilation is NONDETERMINISTIC in a way that is invisible until a run fails, and four
6
6
  // failure modes showed up repeatedly against a real suite:
7
7
  //
8
8
  // 1. An explicit directive silently vanishes. The same prose ("Launch the app WITH ITS
@@ -18,8 +18,12 @@
18
18
  // against later builds. A test exercising none of its subject reporting success is the
19
19
  // worst failure mode a testing tool has, which is why the two rules that detect it are
20
20
  // the only FATAL findings here.
21
+ // 4. A selector names an id the test never gave (issue #148). The compiler cannot see the
22
+ // app, so it composed one from the app's naming pattern, and 3 of 4 such guesses named
23
+ // nothing in the app. A wait/assert is never repaired, so the test went red — and each
24
+ // recompile guessed differently, so it flaked.
21
25
  //
22
- // 1 and 2 are cheap to detect and cheap to fix: hand the finding back to the model and let
26
+ // 1, 2 and 4 are cheap to detect and cheap to fix: hand the finding back to the model and let
23
27
  // it compile once more. That is far better than the alternative, which is a plan that is
24
28
  // quietly wrong and burns a device run to say so. 3 gets the same guided recompile, but a
25
29
  // plan that STILL does not cover its test must never run — see `fatal` below.
@@ -34,9 +38,11 @@ exports.instructionLines = instructionLines;
34
38
  exports.instructionUnits = instructionUnits;
35
39
  exports.actionNodes = actionNodes;
36
40
  exports.tailAnchors = tailAnchors;
41
+ exports.ungroundedIds = ungroundedIds;
37
42
  exports.lintPlan = lintPlan;
38
43
  exports.looksTruncated = looksTruncated;
39
44
  const ir_1 = require("./ir");
45
+ const selector_1 = require("../ui/selector");
40
46
  /** Walk every node in the plan, including control-node bodies. */
41
47
  function* walk(nodes) {
42
48
  for (const n of nodes) {
@@ -257,15 +263,95 @@ function planMentions(tokens, anchor) {
257
263
  return b.length >= 3 && (b.includes(a) || a.includes(b));
258
264
  });
259
265
  }
266
+ // --- ids: does every id the plan selects by come from the prose? (issue #148) ------------
267
+ /** Commands whose FIRST positional is a selector. The rest of `text`'s positionals, and every
268
+ * positional of `type`/`launch`/…, is data — `type @handle` types a handle, it selects nothing. */
269
+ const SELECTOR_FIRST = new Set(['tap', 'click', 'text', 'wait', 'assert']);
270
+ /** Every string the engine resolves as a selector — unlike `planTokens`, which also returns data. */
271
+ function selectorsOf(plan) {
272
+ const out = [];
273
+ for (const n of walk(plan.steps)) {
274
+ if (n.type === 'command') {
275
+ if (SELECTOR_FIRST.has(n.command) && n.positionals[0])
276
+ out.push(n.positionals[0]);
277
+ if (n.command === 'swipe' || n.command === 'scroll') {
278
+ for (const f of n.flags)
279
+ if (f.name === 'on')
280
+ out.push(f.value);
281
+ }
282
+ }
283
+ else if (n.type === 'when') {
284
+ out.push(...n.branches.map((b) => b.selector));
285
+ }
286
+ else {
287
+ out.push(n.selector);
288
+ }
289
+ }
290
+ return out;
291
+ }
292
+ /** The id values (`@x` / `id:x`) the plan selects by, parsed by the engine's own parser so the
293
+ * two can never disagree about what an id selector is — trailing state modifiers included. */
294
+ function planIds(plan) {
295
+ const out = [];
296
+ for (const raw of selectorsOf(plan)) {
297
+ let sel;
298
+ try {
299
+ sel = (0, selector_1.parseSelector)(raw);
300
+ }
301
+ catch {
302
+ continue; // an empty value or contradictory modifiers: a runtime error, not a guess
303
+ }
304
+ if (sel.kind === 'id' && !out.includes(sel.value))
305
+ out.push(sel.value);
306
+ }
307
+ return out;
308
+ }
309
+ /** `com.app:id/login` and `login` name the same element. */
310
+ const RESOURCE_PKG_RE = /^[\w.]+:id\//;
311
+ /** `{{ctx.i}}` / `{{env.X}}` are filled at run time, so they are never part of a guess. */
312
+ const PLACEHOLDER_RE = /\{\{[^}]*\}\}/;
313
+ /** The literal parts of an id worth judging. Under three characters a substring matches
314
+ * anything, so a part that short has no opinion. */
315
+ const idParts = (id) => id
316
+ .replace(RESOURCE_PKG_RE, '')
317
+ .split(PLACEHOLDER_RE)
318
+ .map(normToken)
319
+ .filter((p) => p.length >= 3);
320
+ /** Every word of the prose, and every quoted span whole — an iOS identifier may hold spaces. */
321
+ function proseTokens(nl) {
322
+ const out = [];
323
+ for (const m of nl.matchAll(ANCHOR_RE))
324
+ out.push(m[1] ?? m[2] ?? m[3] ?? '');
325
+ out.push(...nl.split(/[\s`"“”'‘’()[\]{}<>,;!?]+/));
326
+ return out.map(normToken).filter((t) => t.length > 0);
327
+ }
328
+ /**
329
+ * The ids the plan selects by that neither the prose nor `seed` gives — each one a guess.
330
+ *
331
+ * LENIENT on purpose, since a false positive costs a recompile: an id is grounded when every
332
+ * literal part of it (package qualifier and placeholders dropped, case and punctuation ignored)
333
+ * sits inside ONE token of the prose, so `@spinner` is grounded by `vk_spinner` and
334
+ * `id:option_{{ctx.i}}` by `option_0`. One token, never two: `settings_tab` built from the words
335
+ * "Settings tab" is precisely how a guess is made.
336
+ *
337
+ * `seed` is a prior plan the compiler was told to reuse. A green run re-persists the HEALED plan
338
+ * and a repair takes its id off a live screen, so an id a seed already uses is grounded too —
339
+ * pass one only where it can hold a repair (see `compileFromSegments` in ../cli.ts).
340
+ */
341
+ function ungroundedIds(nl, plan, seed) {
342
+ const ground = [...proseTokens(nl), ...(seed ? planIds(seed).flatMap(idParts) : [])];
343
+ return planIds(plan).filter((id) => idParts(id).some((part) => !ground.some((t) => t.includes(part))));
344
+ }
260
345
  /**
261
346
  * Check a compiled plan against the prose it came from.
262
347
  *
263
348
  * @param nl the natural-language test, verbatim
264
349
  * @param plan the plan the model just produced
350
+ * @param seed the prior plan the model was handed to adapt, if any — its ids are not guesses
265
351
  * @returns findings; empty means the plan is consistent with the prose. A finding with
266
352
  * `fatal` set means the plan does not cover the test and must not be run.
267
353
  */
268
- function lintPlan(nl, plan) {
354
+ function lintPlan(nl, plan, seed) {
269
355
  const findings = [];
270
356
  if (FRESH_START_RE.test(nl) && !hasLeafWithFlag(plan, 'launch', 'clear')) {
271
357
  findings.push({
@@ -281,6 +367,21 @@ function lintPlan(nl, plan) {
281
367
  '(skip when absent), or behind when (when the screen is one of several known kinds).',
282
368
  });
283
369
  }
370
+ // Not fatal: a guess that happens to be right still runs, and a wrong one fails its step
371
+ // loudly, so the guided recompile is the remedy. The exception is an ABSENCE check (`--gone`,
372
+ // a guard that skips): it passes on any selector that never matches, a wrong label as much as
373
+ // a guessed id, so making this rule fatal would not close it.
374
+ const guessed = ungroundedIds(nl, plan, seed);
375
+ if (guessed.length > 0) {
376
+ const named = guessed.slice(0, 5).map((id) => JSON.stringify(id)).join(', ');
377
+ findings.push({
378
+ message: `The plan selects by id ${named}${guessed.length > 5 ? ` and ${guessed.length - 5} more` : ''}, but the test ` +
379
+ `never gives ${guessed.length === 1 ? 'that id' : 'those ids'}. You cannot see the app, so an id the test does not state is a guess — it matches ` +
380
+ 'nothing, and a wait/assert on it fails the test with no repair. Use an id only when the test writes it; ' +
381
+ 'select an element the test names by its label with text:<label>, and do not check something the test ' +
382
+ 'names nothing selectable for.',
383
+ });
384
+ }
284
385
  if (!coverageChecksEnabled())
285
386
  return findings;
286
387
  const units = instructionUnits(nl);
@@ -17,18 +17,21 @@ exports.remoteDeviceList = remoteDeviceList;
17
17
  exports.createRemoteBackend = createRemoteBackend;
18
18
  const node_fs_1 = require("node:fs");
19
19
  const node_crypto_1 = require("node:crypto");
20
+ const node_http_1 = require("node:http");
21
+ const node_https_1 = require("node:https");
20
22
  const node_path_1 = require("node:path");
21
23
  const errors_1 = require("../errors");
24
+ const install_timeouts_1 = require("../install-timeouts");
22
25
  const output_1 = require("../output");
23
26
  const rpc_1 = require("../rpc");
24
27
  /**
25
- * The ceiling NONE of the per-call timeouts below can exceed, whatever they say.
28
+ * The ceiling fetch-based calls below cannot exceed, whatever their own budgets say.
26
29
  *
27
30
  * Node's global `fetch` is undici, whose `headersTimeout` and `bodyTimeout` both default to
28
31
  * 300s, and there is no dependency-free way to raise them: a `dispatcher` needs `undici`
29
- * itself, which is bundled but not importable. The `AbortController` below is therefore a
30
- * FLOOR on how long a call may take, never a ceiling — `EXEC_TIMEOUT_MS` says 600s and gets
31
- * 300s.
32
+ * itself, which is bundled but not importable. Long-running install uploads therefore use
33
+ * Node's built-in http/https client instead; its explicit 15-minute timer is the real
34
+ * upload-and-response ceiling. The other AbortControllers remain a floor when set above 300s.
32
35
  *
33
36
  * MEASURED on Node v20.20.2 against a server that held its headers for 310s: the fetch
34
37
  * rejected at 301s with `TypeError: fetch failed`, cause `HeadersTimeoutError`, code
@@ -39,11 +42,10 @@ const rpc_1 = require("../rpc");
39
42
  const FETCH_HEADERS_CEILING_MS = 300_000;
40
43
  // Per-call ceilings. exec is generous: a single leaf may legitimately block for its
41
44
  // whole auto-wait window or an explicit `wait --timeout`, plus device time. Anything here
42
- // above FETCH_HEADERS_CEILING_MS is aspirational — see that constant.
45
+ // above FETCH_HEADERS_CEILING_MS is aspirational unless it uses requestWithNodeHttp.
43
46
  const HEALTH_TIMEOUT_MS = 10_000;
44
47
  const ELEMENTS_TIMEOUT_MS = 60_000;
45
48
  const EXEC_TIMEOUT_MS = 10 * 60_000;
46
- const INSTALL_TIMEOUT_MS = 15 * 60_000;
47
49
  const DEVICE_LIST_TIMEOUT_MS = 30_000;
48
50
  // Meant to sit above the server's own 4-minute boot ceiling, so the SERVER reports why a
49
51
  // boot timed out rather than the client aborting first. It is exactly AT
@@ -52,6 +54,54 @@ const DEVICE_LIST_TIMEOUT_MS = 30_000;
52
54
  const DEVICE_START_TIMEOUT_MS = 5 * 60_000;
53
55
  const DEVICE_STOP_TIMEOUT_MS = 60_000;
54
56
  const trimUrl = (url) => url.replace(/\/+$/, '');
57
+ /**
58
+ * A dependency-free request path whose caller-owned timeout covers BOTH upload and response.
59
+ *
60
+ * Global fetch cannot wait past undici's fixed 300s header/body ceilings. An install may
61
+ * legitimately need 15 minutes, so it uses Node's native protocol clients instead of claiming
62
+ * a budget the transport cannot honor. Buffering the response preserves the Response boundary
63
+ * used below; install responses are small JSON descriptors, never artifacts.
64
+ */
65
+ function requestWithNodeHttp(url, method, headers, body, timeoutMs) {
66
+ const endpoint = new URL(url);
67
+ const send = endpoint.protocol === 'http:' ? node_http_1.request : endpoint.protocol === 'https:' ? node_https_1.request : null;
68
+ if (!send)
69
+ return Promise.reject(new Error(`unsupported server protocol '${endpoint.protocol}'`));
70
+ const payload = body === undefined ? undefined : Buffer.isBuffer(body) ? body : Buffer.from(body);
71
+ const requestHeaders = {
72
+ ...headers,
73
+ ...(payload ? { 'content-length': String(payload.byteLength) } : {}),
74
+ };
75
+ return new Promise((resolve, reject) => {
76
+ let settled = false;
77
+ let timer;
78
+ const finish = (fn) => {
79
+ if (settled)
80
+ return;
81
+ settled = true;
82
+ if (timer)
83
+ clearTimeout(timer);
84
+ fn();
85
+ };
86
+ const req = send(endpoint, { method, headers: requestHeaders }, (res) => {
87
+ const chunks = [];
88
+ res.on('data', (chunk) => chunks.push(Buffer.from(chunk)));
89
+ res.on('error', (e) => finish(() => reject(e)));
90
+ res.on('aborted', () => finish(() => reject(new Error('the server aborted its response'))));
91
+ res.on('end', () => {
92
+ const status = res.statusCode ?? 500;
93
+ finish(() => resolve(new Response(Buffer.concat(chunks), { status })));
94
+ });
95
+ });
96
+ req.on('error', (e) => finish(() => reject(e)));
97
+ timer = setTimeout(() => {
98
+ const timeout = Object.assign(new Error('This operation was aborted'), { name: 'AbortError' });
99
+ req.destroy(timeout);
100
+ }, timeoutMs);
101
+ timer.unref?.();
102
+ req.end(payload);
103
+ });
104
+ }
55
105
  /**
56
106
  * Turn a non-2xx into the error the caller sees.
57
107
  *
@@ -132,24 +182,31 @@ class RemoteTransport {
132
182
  h.authorization = `Bearer ${this.opts.authKey}`;
133
183
  return h;
134
184
  }
135
- async request(method, path, body, timeoutMs, extraHeaders = {}) {
185
+ async request(method, path, body, timeoutMs, extraHeaders = {}, transport = 'fetch') {
136
186
  const url = `${this.base}${path}`;
137
- const controller = new AbortController();
138
- const timer = setTimeout(() => controller.abort(), timeoutMs);
187
+ let timer;
139
188
  let res;
140
189
  try {
141
- res = await fetch(url, {
142
- method,
143
- headers: this.headers(extraHeaders),
144
- body,
145
- signal: controller.signal,
146
- });
190
+ if (transport === 'node-http') {
191
+ res = await requestWithNodeHttp(url, method, this.headers(extraHeaders), body, timeoutMs);
192
+ }
193
+ else {
194
+ const controller = new AbortController();
195
+ timer = setTimeout(() => controller.abort(), timeoutMs);
196
+ res = await fetch(url, {
197
+ method,
198
+ headers: this.headers(extraHeaders),
199
+ body,
200
+ signal: controller.signal,
201
+ });
202
+ }
147
203
  }
148
204
  catch (e) {
149
205
  throw new errors_1.CliError(`cannot reach verikun server at ${url} (${transportReason(e, timeoutMs)})`, 3);
150
206
  }
151
207
  finally {
152
- clearTimeout(timer);
208
+ if (timer)
209
+ clearTimeout(timer);
153
210
  }
154
211
  if (!res.ok) {
155
212
  const body = await readBody(res);
@@ -262,11 +319,11 @@ function createRemoteBackend(opts, health) {
262
319
  throw new errors_1.CliError(`install: cannot read '${appPath}' (${e.message})`, 2);
263
320
  }
264
321
  const sha256 = (0, node_crypto_1.createHash)('sha256').update(buf).digest('hex');
265
- const res = await t.request('POST', '/v1/install', buf, INSTALL_TIMEOUT_MS, {
322
+ const res = await t.request('POST', '/v1/install', buf, install_timeouts_1.REMOTE_INSTALL_TIMEOUT_MS, {
266
323
  'content-type': 'application/octet-stream',
267
324
  'x-verikun-ext': ext,
268
325
  'x-verikun-sha256': sha256,
269
- });
326
+ }, 'node-http');
270
327
  // Install is the one operation the server replays elsewhere, so a move here means
271
328
  // the build DID land — on a different device than the one we started with.
272
329
  if (res.deviceChanged)
package/dist/cli.js CHANGED
@@ -1905,7 +1905,14 @@ async function compileFromSegments(segments, key, opts, cost, provider) {
1905
1905
  // FREE, and refusing it over a ceiling we were never about to spend against would fail
1906
1906
  // a test for somebody else's tokens.
1907
1907
  assertBudgetForCompile(cost, opts.maxCostUsd, `compiling ${where} — the test is only partly compiled`);
1908
- const seed = seedPlan(segKey, where);
1908
+ // A section's entry is only ever raw compile output — a green run re-persists the WHOLE
1909
+ // test's key, never a section's — so an id in it that its own prose never gives can only be
1910
+ // a guess (issue #148), and handing it back as "reuse this" would keep it alive.
1911
+ let seed = seedPlan(segKey, where);
1912
+ if (seed && (0, lint_1.ungroundedIds)(seg.text, seed.plan).length > 0) {
1913
+ (0, output_1.err)(`[ai] ${where}: ignoring a prior plan that guesses ids its prose never gives — compiling fresh`);
1914
+ seed = null;
1915
+ }
1909
1916
  (0, output_1.err)(`[ai] ${where}: compiling with ${opts.model}…${lock.degraded ? ` (${lock.degraded})` : ''}`);
1910
1917
  let compiled;
1911
1918
  try {
@@ -1939,6 +1946,14 @@ async function compileFromSegments(segments, key, opts, cost, provider) {
1939
1946
  (0, output_1.err)(`[ai] ${where}: the compiled section does not cover its prose (${compiled.plan.steps.length} step(s)) — not caching it; compiling the test as one instead`);
1940
1947
  return null; // the finally below releases the lock
1941
1948
  }
1949
+ // A guessed id in a FRAGMENT is issue #148 spliced into every test that includes it. Same
1950
+ // answer as a short section: keep it out of the cache and compile the test whole, where the
1951
+ // lint hands the guess back for one guided recompile.
1952
+ const guessed = (0, lint_1.ungroundedIds)(seg.text, compiled.plan);
1953
+ if (guessed.length > 0) {
1954
+ (0, output_1.err)(`[ai] ${where}: the compiled section selects by ids its prose never gives (${guessed.join(', ')}) — not caching it; compiling the test as one instead`);
1955
+ return null; // the finally below releases the lock
1956
+ }
1942
1957
  // INSIDE the lock and before the release: a waiter re-reads the cache the instant the
1943
1958
  // lock disappears, so releasing first would hand it a miss and buy the second compile
1944
1959
  // this whole mechanism exists to prevent.
@@ -2013,7 +2028,10 @@ async function obtainPlan(key, file, opts, cost, provider, segments = []) {
2013
2028
  // sometimes only the test's opening compiles at all. One guided retry is much cheaper
2014
2029
  // than discovering it as a device-run failure several steps later, or (worse, for a
2015
2030
  // truncation) as a pass. Budget is re-checked HERE: the first attempt is already billed.
2016
- let remaining = (0, lint_1.lintPlan)(key.nl, compiled.plan);
2031
+ //
2032
+ // The seed goes in too: a green run re-persists the HEALED plan, whose repaired ids came off a
2033
+ // live screen, so flagging them would push every new build off a selector the device confirmed.
2034
+ let remaining = (0, lint_1.lintPlan)(key.nl, compiled.plan, seed?.plan);
2017
2035
  if (remaining.length > 0) {
2018
2036
  const feedback = remaining.map((f) => `- ${f.message}`).join('\n');
2019
2037
  (0, output_1.err)(`[ai] compiled plan does not match the test — recompiling once:\n${feedback}`);
@@ -2029,7 +2047,7 @@ async function obtainPlan(key, file, opts, cost, provider, segments = []) {
2029
2047
  retryFeedback: feedback,
2030
2048
  });
2031
2049
  cost.add(retry.usage, 'compile');
2032
- const still = (0, lint_1.lintPlan)(key.nl, retry.plan);
2050
+ const still = (0, lint_1.lintPlan)(key.nl, retry.plan, seed?.plan);
2033
2051
  // Keep the retry either way: it was compiled with strictly more information. If it
2034
2052
  // still trips the lint, say so rather than pretending the plan is clean — and for a
2035
2053
  // coverage finding, do not claim it will run, because the gate below rejects it.
@@ -33,6 +33,7 @@ const ARTIFACT_RULES = [
33
33
  const TRANSIENT_RULES = [[/INSTALL_FAILED_ABORTED/, 'the install session was aborted']];
34
34
  // --- fast paths: name the reason and skip the probe. NEVER the gate. ----------
35
35
  const UNREACHABLE_RULES = [
36
+ [/stopped responding.*install/i, 'the device stopped responding during install'],
36
37
  [/device (?:'[^']*' )?not found/i, 'the device is not attached'],
37
38
  [/no devices\/emulators found/i, 'no device is attached'],
38
39
  [/device offline/i, 'the device is offline'],
@@ -49,6 +50,9 @@ const DEVICE_STATE_RULES = [
49
50
  // `install` already tried removing it and reinstalling (drivers/adb.ts); reaching here
50
51
  // means that did not work, so the wording must not imply nothing was attempted.
51
52
  [/INSTALL_FAILED_UPDATE_INCOMPATIBLE/, 'a differently-signed build could not be replaced on the device'],
53
+ // `install` already retries with `-d` (drivers/adb.ts); reaching here means the build
54
+ // or device isn't debuggable, so Android is still refusing — a genuine per-device
55
+ // conflict, not a sign the retry never ran.
52
56
  [/INSTALL_FAILED_VERSION_DOWNGRADE/, 'the device holds a newer build of this package'],
53
57
  [/INSTALL_FAILED_ALREADY_EXISTS/, 'the package is already installed on the device'],
54
58
  [/INSTALL_FAILED_DUPLICATE_PERMISSION/, 'another app on the device declares one of these permissions'],
@@ -166,8 +170,9 @@ function classifyInstallFailure(e) {
166
170
  if (named)
167
171
  return { move: true, kind: 'device-state', reason: named };
168
172
  // The inversion. Not in the denylist ⇒ it is about the device, even though we have
169
- // never seen this wording. Bounded by MAX_FAILOVER_HOPS, and the caller reports the
170
- // FIRST device's error on exhaustion, so a wrong guess costs time, not diagnosis.
173
+ // never seen this wording. Bounded by the request-derived install hop limit, and the
174
+ // caller reports the FIRST device's error on exhaustion, so a wrong guess costs time,
175
+ // not diagnosis.
171
176
  return { move: true, kind: 'device-state', reason: 'the device could not install this build', unclassified: true };
172
177
  }
173
178
  return classify(e, UNKNOWN_STAY);
@@ -1104,21 +1104,27 @@ class AdbDriver {
1104
1104
  throw new errors_1.CliError(signatureConflictHelp(appPath, pkg, 'retry-failed', second), 3);
1105
1105
  }
1106
1106
  /**
1107
- * One `adb install -r` attempt: null on success, else adb's collapsed output.
1107
+ * One `adb install -r -d` attempt: null on success, else adb's collapsed output.
1108
1108
  *
1109
1109
  * Returns rather than throws because the caller has to READ the failure to decide
1110
1110
  * whether it is recoverable — and a thrown-and-caught CliError would put the final
1111
1111
  * message's own prefix inside the message it is composing.
1112
1112
  *
1113
1113
  * `-r` reinstalls over an existing package keeping its data (the common
1114
- * update-the-build-under-test case). A large APK can legitimately take minutes to
1115
- * stream + install, so the timeout is far above the 30s default. adb reports failures
1116
- * both as a non-zero exit AND as a `Failure [REASON]` line on stdout with exit 0
1117
- * (varies by adb version) — check both, and require `Success` positively rather than
1118
- * merely inferring it from the absence of `Failure`.
1114
+ * update-the-build-under-test case). `-d` allows a lower versionCode to replace a
1115
+ * higher one already on the device — unconditional here, not a retry like the
1116
+ * signature-conflict replace below, because it carries no data-loss trade-off (inert
1117
+ * when there is no downgrade, and keeps data the same way `-r` does when there is).
1118
+ * Android only honors it for a debuggable build, so a release-signed downgrade still
1119
+ * surfaces `INSTALL_FAILED_VERSION_DOWNGRADE` unresolved — see the comment on that
1120
+ * code in `device/failover.ts`. A large APK can legitimately take minutes to stream +
1121
+ * install, so the timeout is far above the 30s default. adb reports failures both as a
1122
+ * non-zero exit AND as a `Failure [REASON]` line on stdout with exit 0 (varies by adb
1123
+ * version) — check both, and require `Success` positively rather than merely inferring
1124
+ * it from the absence of `Failure`.
1119
1125
  */
1120
1126
  tryInstall(appPath) {
1121
- const r = (0, exec_1.runText)(ADB, this.withSerial(['install', '-r', appPath]), { timeout: 10 * 60 * 1000 });
1127
+ const r = (0, exec_1.runText)(ADB, this.withSerial(['install', '-r', '-d', appPath]), { timeout: 10 * 60 * 1000 });
1122
1128
  const combined = `${r.stdout}\n${r.stderr}`;
1123
1129
  if (r.code !== 0 || /^Failure\b/im.test(combined) || !/^Success\b/im.test(combined)) {
1124
1130
  return combined.replace(/\s+/g, ' ').trim() || `exit code ${r.code}`;
@@ -0,0 +1,12 @@
1
+ "use strict";
2
+ // The remote-install time budget is shared by the client that owns the request and the
3
+ // server that may retry it. Keeping the numbers together prevents a server-side retry plan
4
+ // that can never finish before the client gives up.
5
+ Object.defineProperty(exports, "__esModule", { value: true });
6
+ exports.MAX_INSTALL_FAILOVER_HOPS = exports.INSTALL_DEVICE_TIMEOUT_MS = exports.REMOTE_INSTALL_TIMEOUT_MS = void 0;
7
+ /** Whole upload + install request, enforced by the remote client. */
8
+ exports.REMOTE_INSTALL_TIMEOUT_MS = 15 * 60_000;
9
+ /** One device gets four minutes; a measured ~170s large install still has useful margin. */
10
+ exports.INSTALL_DEVICE_TIMEOUT_MS = 4 * 60_000;
11
+ /** Three attempts consume at most 12m, leaving 3m of the request budget for the upload. */
12
+ exports.MAX_INSTALL_FAILOVER_HOPS = Math.max(0, Math.floor(exports.REMOTE_INSTALL_TIMEOUT_MS / exports.INSTALL_DEVICE_TIMEOUT_MS) - 1);
package/dist/server.js CHANGED
@@ -45,6 +45,7 @@ const node_stream_1 = require("node:stream");
45
45
  const promises_1 = require("node:stream/promises");
46
46
  const args_1 = require("./args");
47
47
  const errors_1 = require("./errors");
48
+ const install_timeouts_1 = require("./install-timeouts");
48
49
  const drivers_1 = require("./drivers");
49
50
  const manager_1 = require("./companion/manager");
50
51
  const claims_1 = require("./device/claims");
@@ -86,10 +87,6 @@ const INSTALL_BODY_CAP = 512 * 1024 * 1024; // 512 MB app build
86
87
  // to survive a client-side compile/repair pause, short enough that a crashed
87
88
  // caller doesn't wedge the device.
88
89
  const LOCK_IDLE_MS = 5 * 60 * 1000;
89
- // How many times ONE request may move device. 2 moves = 3 devices tried, which sits
90
- // comfortably inside the client's 15-minute install ceiling at ~1 minute an install,
91
- // while a farm of ten wedged emulators cannot burn ten installs inside one request.
92
- const MAX_FAILOVER_HOPS = 2;
93
90
  // How often to ask whether the host's adb server has rotted. Generous on purpose: the
94
91
  // check shells out to `log show` (~1s) and the condition it looks for accumulates over
95
92
  // DAYS, so a tight interval would buy nothing and spend host time on every idle server.
@@ -108,6 +105,16 @@ const ADB_RECYCLE_CHECK_MS = 10 * 60 * 1000;
108
105
  */
109
106
  const RECONCILE_INTERVAL_MS = 60_000;
110
107
  const PROBE_RETRY_MS = 1000;
108
+ /** Kept distinct so an all-device timeout is never mistaken for evidence of a bad build. */
109
+ class InstallAttemptTimeoutError extends errors_1.CliError {
110
+ constructor(serial, timeoutMs) {
111
+ const duration = timeoutMs >= 1000 ? `${Math.round(timeoutMs / 1000)}s` : `${timeoutMs}ms`;
112
+ super(`device ${serial} stopped responding: install exceeded its ${duration} per-device deadline`, 3);
113
+ this.name = 'InstallAttemptTimeoutError';
114
+ }
115
+ }
116
+ /** Failover wraps an exhausted attempt in HttpError, so retain a wire-safe marker too. */
117
+ const isInstallAttemptTimeout = (e) => e instanceof InstallAttemptTimeoutError || (e instanceof Error && /per-device deadline/.test(e.message));
111
118
  // Deliberately below the client's 5-minute ceiling, so a slow boot is reported by the
112
119
  // side that knows WHY ("did not finish booting within 240s") rather than as a generic
113
120
  // client-side abort.
@@ -1251,7 +1258,7 @@ function buildServer(config) {
1251
1258
  * On exhaustion it throws the FIRST device's error, never the last. That inversion is
1252
1259
  * what makes move-by-default safe: a wrong move costs time, not the diagnosis.
1253
1260
  */
1254
- async function installWithFailover(serial, tmpPath) {
1261
+ async function installWithFailover(serial, tmpPath, timedOut) {
1255
1262
  let change;
1256
1263
  let moves = 0;
1257
1264
  let firstError;
@@ -1291,7 +1298,7 @@ function buildServer(config) {
1291
1298
  // are still on our disk), so it hops rather than throwing a bare 503 that never
1292
1299
  // reaches the failover machinery at all.
1293
1300
  const gone = firstError ?? new server_http_1.HttpError(503, `device ${from} is no longer attached`, 3);
1294
- if (!config.failover || hop >= MAX_FAILOVER_HOPS)
1301
+ if (!config.failover || hop >= install_timeouts_1.MAX_INSTALL_FAILOVER_HOPS)
1295
1302
  throw gone;
1296
1303
  // `unreachable` is the literal truth — it is not in the pool — and it is also
1297
1304
  // inert here: `shrink` sees a non-member and takes its cleanup tail either way.
@@ -1299,7 +1306,28 @@ function buildServer(config) {
1299
1306
  continue;
1300
1307
  }
1301
1308
  try {
1302
- await handle.install(tmpPath);
1309
+ const timeoutMs = config.installAttemptMs ?? install_timeouts_1.INSTALL_DEVICE_TIMEOUT_MS;
1310
+ await new Promise((resolve, reject) => {
1311
+ let settled = false;
1312
+ const finish = (fn) => {
1313
+ if (settled)
1314
+ return;
1315
+ settled = true;
1316
+ clearTimeout(timer);
1317
+ fn();
1318
+ };
1319
+ const timer = setTimeout(() => {
1320
+ const timeout = new InstallAttemptTimeoutError(from, timeoutMs);
1321
+ // A race alone would leave the worker blocked in adb forever. Retiring it is
1322
+ // what releases the transport and makes the deadline a recovery mechanism.
1323
+ timedOut.add(from);
1324
+ quarantineDevice(from, timeout.message);
1325
+ pool.retire(from);
1326
+ finish(() => reject(timeout));
1327
+ }, timeoutMs);
1328
+ timer.unref?.();
1329
+ void handle.install(tmpPath).then(() => finish(resolve), (e) => finish(() => reject(e)));
1330
+ });
1303
1331
  return { change, moves };
1304
1332
  }
1305
1333
  catch (e) {
@@ -1310,8 +1338,18 @@ function buildServer(config) {
1310
1338
  noteVerdict(verdict, e, 'install');
1311
1339
  // The artifact is broken / the caller is wrong / failover is off / we are out of
1312
1340
  // hops: report the first failure unchanged, exactly as before this feature.
1313
- if (!verdict.move || !config.failover || hop >= MAX_FAILOVER_HOPS)
1341
+ if (!verdict.move || !config.failover || hop >= install_timeouts_1.MAX_INSTALL_FAILOVER_HOPS) {
1342
+ // A timeout retired its worker before this catch. If it cannot move, finish the
1343
+ // cleanup that pickFailoverDevice would otherwise own; especially, never leave
1344
+ // a host-global claim pinned to a worker that no longer exists.
1345
+ if (isInstallAttemptTimeout(e)) {
1346
+ evictHoldersOf(from, `${from} left the pool after its install timed out`);
1347
+ (0, manager_1.releaseCompanionOn)(from);
1348
+ if ((0, claims_1.claimsEnabled)(claimEnv))
1349
+ (0, claims_1.releaseClaim)(from, { ...claimOpts, mineOnly: true });
1350
+ }
1314
1351
  throw firstError;
1352
+ }
1315
1353
  from = await hopOrThrow(verdict.reason, firstError, verdict.kind);
1316
1354
  }
1317
1355
  }
@@ -1361,9 +1399,10 @@ function buildServer(config) {
1361
1399
  throw new server_http_1.HttpError(503, `no device is left to install onto${lostDevice ? ` — last loss: ${lostDevice}` : ''}`, 3);
1362
1400
  }
1363
1401
  (0, output_1.err)(`[server] install: received ${size} bytes (.${ext}), installing on ${targets.join(', ')}…`);
1402
+ const timedOut = new Set();
1364
1403
  const outcomes = await Promise.all(targets.map(async (serial) => {
1365
1404
  try {
1366
- return { serial, ...(await installWithFailover(serial, tmpPath)), error: null };
1405
+ return { serial, ...(await installWithFailover(serial, tmpPath, timedOut)), error: null };
1367
1406
  }
1368
1407
  catch (e) {
1369
1408
  return { serial, change: undefined, moves: 0, error: e };
@@ -1387,9 +1426,21 @@ function buildServer(config) {
1387
1426
  // polarity, since the device side is open-ended and OEM-specific. That is right
1388
1427
  // for ONE device failing; applied to every device at once it condemns the whole
1389
1428
  // pool for what this very branch has just concluded is a bad build. Undo them.
1390
- const condemned = targets.filter((t) => quarantine.delete(t));
1391
- if (condemned.length) {
1392
- (0, output_1.err)(`[server] install: failed on every device, so the build is the suspect — un-quarantining ${condemned.join(', ')}`);
1429
+ const hasTimeout = timedOut.size > 0;
1430
+ if (!hasTimeout) {
1431
+ // No lane succeeded and none hung: the common input is now the stronger
1432
+ // suspect, so undo the per-device quarantines made by the move-by-default
1433
+ // classifier.
1434
+ const condemned = targets.filter((t) => quarantine.delete(t));
1435
+ if (condemned.length) {
1436
+ (0, output_1.err)(`[server] install: failed on every device, so the build is the suspect — un-quarantining ${condemned.join(', ')}`);
1437
+ }
1438
+ }
1439
+ else {
1440
+ // A deadline is direct evidence about a device, not the artifact. Preserve
1441
+ // every quarantine when any lane timed out; otherwise the next request could
1442
+ // immediately be dealt the same wedged worker again.
1443
+ (0, output_1.err)('[server] install: every device failed and at least one timed out — keeping the device quarantines');
1393
1444
  }
1394
1445
  throw failed[0].error;
1395
1446
  }
@@ -1412,11 +1463,15 @@ function buildServer(config) {
1412
1463
  skipped = failed.map((f) => {
1413
1464
  const serial = landedOn(f);
1414
1465
  const reason = (0, server_http_1.firstLine)(f.error.message);
1415
- if (pool.serials().includes(serial)) {
1466
+ const serving = pool.serials().includes(serial);
1467
+ if (serving)
1416
1468
  pool.retire(serial);
1417
- // Same two rules as the shed in `shrink`: `degraded` is for MEMBERS, and a
1418
- // device that is not serving belongs in `quarantine` — which `rejoinDevice`
1419
- // clears on the evidence of a worker that started and a build that installed.
1469
+ // Same two rules as the shed in `shrink`: `degraded` is for MEMBERS, and a
1470
+ // device that is not serving belongs in `quarantine` — which `rejoinDevice`
1471
+ // clears on the evidence of a worker that started and a build that installed.
1472
+ // The timeout path retired its worker already, but still needs this common
1473
+ // bookkeeping tail. Other already-absent failures keep their existing handling.
1474
+ if (serving || isInstallAttemptTimeout(f.error)) {
1420
1475
  degraded.delete(serial);
1421
1476
  quarantineDevice(serial, `did not take the current build — ${reason}`);
1422
1477
  evictHoldersOf(serial, `${serial} left the pool without the current build`);
package/dist/version.js CHANGED
@@ -3,4 +3,4 @@ Object.defineProperty(exports, "__esModule", { value: true });
3
3
  exports.VERSION = void 0;
4
4
  // GENERATED by scripts/gen-version.mjs from package.json's "version" at build time
5
5
  // (the `prebuild` script). Do NOT edit by hand; bump package.json instead.
6
- exports.VERSION = '0.28.0';
6
+ exports.VERSION = '0.29.1';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "verikun",
3
- "version": "0.28.0",
3
+ "version": "0.29.1",
4
4
  "description": "Drive Android emulators/devices and iOS simulators for AI agents: tap, type, swipe, screenshot, and inspect the UI hierarchy by semantic identifiers — like Puppeteer for native apps.",
5
5
  "keywords": [
6
6
  "android",