verikun 0.26.0-rc.5 → 0.26.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -390,7 +390,19 @@ vk ai onboarding.md --timeout 5m # tighten the run timeout (default 15m)
390
390
  the flow. An `if-present` guard **waits for its selector to settle** (at least two looks
391
391
  at the screen) before deciding the optional UI is absent, so a dialog that animates in a
392
392
  beat after the transition is still caught. An absent guard costs about one extra UI dump;
393
- `VERIKUN_GUARD_SETTLE_MS=0` restores the old single-shot probe.
393
+ `VERIKUN_GUARD_SETTLE_MS=0` restores the old single-shot probe. A guard that cannot read the
394
+ screen **at all** — the app force-stopped, mid-launch, or busy mid-transition — keeps looking
395
+ for up to 10s rather than aborting; still blind after that is exit `3`, never "absent".
396
+ - **A compile has to cover the test.** Compilation is nondeterministic, and its worst outcome
397
+ is a plan that stops part-way: it asserts nothing after that point, so it *passes*, caches
398
+ green, and replays against later builds — a test exercising none of its subject reporting
399
+ success. A fresh compile is therefore checked against the prose (its size, and whether the
400
+ plan references what the test's closing steps name); a failure buys one guided recompile,
401
+ and a plan still short after that is **rejected at exit `1` and not cached**, so rerunning
402
+ compiles again rather than replaying it. Under `vk suite` that means `--retries` cannot turn
403
+ a truncated compile green. `VERIKUN_NO_COMPILE_CHECK=1` turns the check off. The stderr line
404
+ `compiled N top-level step(s)` (`planSteps` in `--json` and in a suite manifest row) is the
405
+ number to compare across runs of an unchanged test.
394
406
  - **Progress streams to stderr** (never silent in CI); **stdout is the report path**
395
407
  (`--json` for a structured summary). It records like any flow, so it ends with the
396
408
  same JUnit + HTML report — including the token/cost line and **suggested improvements**
@@ -409,7 +421,8 @@ vk ai onboarding.md --timeout 5m # tighten the run timeout (default 15m)
409
421
  ever called to **compile** (once, on a cache miss) or to **repair** (≤3 per failing step);
410
422
  replay is always $0, and every non-`ai` command is $0 always. Full mechanism, the estimate
411
423
  formula and the cache multipliers: <https://ddikman.github.io/verikun/reference/cost/>.
412
- - Exit `0` pass · `1` a step failed (or the budget/timeout was hit) · `2` usage · `3` environment
424
+ - Exit `0` pass · `1` a step failed (or the budget/timeout was hit, or the compile did not
425
+ cover the test) · `2` usage · `3` environment
413
426
  (e.g. the model's API key — `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` — unset, or the `codex` /
414
427
  `cursor-agent` CLI missing / not logged in for `--model codex-cli` / `cursor-cli`).
415
428
  - **`3` also means the device toolchain is broken**, checked *before* anything is compiled
package/CHANGELOG.md CHANGED
@@ -6,43 +6,27 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.26.1] - 2026-09-07
10
+
11
+ Fixes a `vk ai` run dying when the app is redrawing at the moment a guard checks the screen.
12
+
9
13
  ### Added
10
- - **`@include <path>`** in a `vk ai` test inlines another file's prose, so a shared preamble
11
- lives in one place. Paths resolve relative to the including file. ([#112])
12
- - **`vk server --log-file <path|off>`** writes a rotating server log, on by default at
13
- `~/.verikun/logs/server-<port>.log`. `VERIKUN_LOG_FILE` sets the same value.
14
- - **`vk server`** re-adopts pooled devices that come back, sweeping every 60s with backoff.
15
- - **`/v1/health`** reports `degraded`: pool members still serving but dealt last.
16
- - **`VERIKUN_NO_PLAN_LOCK=1`** turns off plan-compile locking, restoring the previous
17
- compile-in-parallel behaviour. ([#117])
14
+ - **`errorKind` on every `vk server` error response**: a device error keeps its class over the
15
+ wire, not only on `/v1/exec`. Absent from older servers. ([#80])
16
+
17
+ ### Fixed
18
+ - **`vk ai` guards** now ride out a transient "no window" for up to 10s instead of aborting the
19
+ run. A bare `vk ui` still exits `3`. ([#80])
18
20
 
19
21
  ### Changed
20
- - **`vk suite`** skips `_`-prefixed `*.md` — those are shared fragments, not tests. ([#112])
21
- - **`vk ai`** compiles each chunk of an `@include`d test separately and caches it, so shared
22
- prose is compiled once across a suite. ([#112])
23
- - **`vk suite`** lanes now wait for a shared `@include` fragment's first compile instead of
24
- each paying for one. ([#117])
25
- - **`vk server`** demotes a failing pooled device instead of dropping it, so capacity no
26
- longer ratchets down and never recovers.
27
- - **`vk server`** deals leases healthy-first, then least-recently-used, instead of first-fit.
28
- - **`vk server`** no longer blames a device when the probe itself fails for a host reason.
29
- - **`vk server`** keeps the pool intact when an install fails on every device — the build is
30
- the suspect, not the phones.
31
- - **`vk server`** logs error bodies, run tokens, leased serials and every lease eviction.
32
- - **`vk install`** (Android) removes a differently-signed build of the same package and retries,
33
- instead of failing. Its app data is lost; stderr says so. ([#96])
34
- - **`VERIKUN_CLAIM_TTL_MIN`** now also paces a parallel suite's claim heartbeat — a quarter of
35
- the window, capped at 60s — so a short TTL no longer races it.
22
+ - **"No window to read"** now names a busy main thread mid-transition as a cause, and no longer
23
+ suggests a command that waits. ([#80])
36
24
 
37
- ### Fixed
38
- - **`vk server`** terminates a worker that stops answering, instead of holding its lease
39
- forever while health still advertises the device.
25
+ [#80]: https://github.com/ddikman/verikun/issues/80
40
26
 
41
- [#96]: https://github.com/ddikman/verikun/issues/96
42
- [#112]: https://github.com/ddikman/verikun/issues/112
43
- [#117]: https://github.com/ddikman/verikun/issues/117
27
+ ## [0.26.0] - 2026-09-05
44
28
 
45
- ## [0.26.0] - 2026-08-24
29
+ Runs a suite across a pool of devices, and shares prose between tests with `@include`.
46
30
 
47
31
  ### Added
48
32
  - **`vk suite --devices a,b` / `--servers u1,u2`**: run tests across a device pool, next free
@@ -61,6 +45,18 @@ All notable changes to this project are documented here. The format is based on
61
45
  working directory don't delete each other's state.
62
46
  - **`POST /v1/lease`** and `capacity` / `devices` on `/v1/health`: which device a run token holds,
63
47
  and how many the server has.
48
+ - **`@include <path>`** in a `vk ai` test inlines another file's prose, so a shared preamble
49
+ lives in one place. Paths resolve relative to the including file. ([#112])
50
+ - **`vk server --log-file <path|off>`** writes a rotating server log, on by default at
51
+ `~/.verikun/logs/server-<port>.log`. `VERIKUN_LOG_FILE` sets the same value.
52
+ - **`vk server`** re-adopts pooled devices that come back, sweeping every 60s with backoff.
53
+ - **`/v1/health`** reports `degraded`: pool members still serving but dealt last.
54
+ - **`VERIKUN_NO_PLAN_LOCK=1`** turns off plan-compile locking, restoring the previous
55
+ compile-in-parallel behaviour. ([#117])
56
+ - **`VERIKUN_NO_COMPILE_CHECK=1`** turns off the `vk ai` compile-coverage check, restoring
57
+ the previous accept-anything behaviour. ([#127])
58
+ - **`planSteps`** in a suite manifest row and `vk ai --json`: the plan's compiled size, next
59
+ to `steps` (what executed). ([#127])
64
60
 
65
61
  ### Changed
66
62
  - **Suite manifest**: adds `totals.wallClockMs`, `concurrency` and a per-test `device`.
@@ -82,13 +78,44 @@ All notable changes to this project are documented here. The format is based on
82
78
  phone unasked.
83
79
  - **`vk server` failover on a pool**: a failed device is quarantined and a healthy one takes its
84
80
  place, so capacity holds. ([#39], [#99])
81
+ - **`vk suite`** skips `_`-prefixed `*.md` — those are shared fragments, not tests. ([#112])
82
+ - **`vk ai`** compiles each chunk of an `@include`d test separately and caches it, so shared
83
+ prose is compiled once across a suite. ([#112])
84
+ - **`vk suite`** lanes now wait for a shared `@include` fragment's first compile instead of
85
+ each paying for one. ([#117])
86
+ - **`vk server`** demotes a failing pooled device instead of dropping it, so capacity no
87
+ longer ratchets down and never recovers.
88
+ - **`vk server`** deals leases healthy-first, then least-recently-used, instead of first-fit.
89
+ - **`vk server`** no longer blames a device when the probe itself fails for a host reason.
90
+ - **`vk server`** keeps the pool intact when an install fails on every device — the build is
91
+ the suspect, not the phones.
92
+ - **`vk server`** logs error bodies, run tokens, leased serials and every lease eviction.
93
+ - **`vk install`** (Android) removes a differently-signed build of the same package and retries,
94
+ instead of failing. Its app data is lost; stderr says so. ([#96])
95
+ - **`VERIKUN_CLAIM_TTL_MIN`** now also paces a parallel suite's claim heartbeat — a quarter of
96
+ the window, capped at 60s — so a short TTL no longer races it.
97
+ - **Contributor docs**: rules stated once, with the history left to the docs site; the companion
98
+ notes no longer describe it as a prototype.
99
+ - **Docs site**: the contracts, core-principles, plan-engine and AI-plans pages state each rule
100
+ once and link the page that owns the detail.
101
+ - **Docs site**: shorter remote-device, troubleshooting and cost guides; global flags now live on
102
+ the Commands page — the standalone page is gone.
85
103
 
86
104
  ### Fixed
87
105
  - **`--server` runs no longer fail with `fetch failed`** when a step follows a pause longer
88
106
  than 5s, such as a cold compile.
89
107
  - **`--server` screenshots and failure evidence** are no longer corrupted in the archived report.
108
+ - **`vk ai`** rejects a compile that covers only the start of its test, instead of caching
109
+ and replaying the short plan as a pass. Exit `1`. ([#127])
110
+ - **`vk ai`** no longer seeds a new compile from a cached plan that is itself truncated. ([#127])
111
+ - **`vk server`** terminates a worker that stops answering, instead of holding its lease
112
+ forever while health still advertises the device.
90
113
 
91
114
  [#39]: https://github.com/ddikman/verikun/issues/39
115
+ [#96]: https://github.com/ddikman/verikun/issues/96
116
+ [#112]: https://github.com/ddikman/verikun/issues/112
117
+ [#117]: https://github.com/ddikman/verikun/issues/117
118
+ [#127]: https://github.com/ddikman/verikun/issues/127
92
119
 
93
120
  ## [0.25.1] - 2026-08-21
94
121
 
package/README.md CHANGED
@@ -6,7 +6,7 @@
6
6
 
7
7
  - **Agent CLI** — `vk <command>`: one-shot commands to inspect the screen as a semantic tree (or screenshot) and act on it.
8
8
  - **Puppeteer for native mobile** — a thin wrapper over native Android and iOS automation runners with zero runtime dependencies.
9
- - **Natural-language tests** — `vk ai <file>`: runs plain-English tests, compiled once and replayed model-free (~$0), calling a model only to self-heal a drifted step. Tests share a preamble with `@include`, written once instead of pasted into each. [What that costs](https://ddikman.github.io/verikun/reference/cost/), and how the `--max-cost-usd` ceiling bounds it.
9
+ - **Natural-language tests** — `vk ai <file>`: runs plain-English tests, compiled once and replayed model-free (~$0), calling a model only to self-heal a drifted step. Tests share a preamble with `@include`, written once instead of pasted into each. A compile that does not cover its test is rejected rather than cached as a pass. [What that costs](https://ddikman.github.io/verikun/reference/cost/), and how the `--max-cost-usd` ceiling bounds it.
10
10
  - **Self-improving** — the agent runner will provide prescriptive improvements to existing scripts to help stabilise flakiness for future runs.
11
11
  - **CI-ready** — `vk suite` runs a folder of tests as one gated pass/fail run, across one device or a whole pool of them; `vk server` exposes real devices over an authenticated tunnel so a disposable CI runner (no phone attached) can still drive them, and routes around any that goes bad.
12
12
 
@@ -63,7 +63,7 @@ There is also no need for an MCP here: verikun runs locally with all its depende
63
63
  |---|---|
64
64
  | **Getting started** | [Installation](https://ddikman.github.io/verikun/getting-started/installation/) · [Your first test](https://ddikman.github.io/verikun/getting-started/your-first-test/) · [Using it from an AI agent](https://ddikman.github.io/verikun/getting-started/using-from-an-agent/) |
65
65
  | **Guides** | [Writing test cases](https://ddikman.github.io/verikun/guides/writing-test-cases/) · [Natural-language tests](https://ddikman.github.io/verikun/guides/natural-language-tests/) · [Suites](https://ddikman.github.io/verikun/guides/suites/) · [Remote devices & CI](https://ddikman.github.io/verikun/guides/remote-devices-and-ci/) · [iOS setup](https://ddikman.github.io/verikun/guides/ios-setup/) · [Platform support](https://ddikman.github.io/verikun/guides/platform-support/) · [Troubleshooting](https://ddikman.github.io/verikun/guides/troubleshooting/) |
66
- | **Reference** | [Commands](https://ddikman.github.io/verikun/reference/commands/) · [Selectors](https://ddikman.github.io/verikun/reference/selectors/) · [Auto-wait](https://ddikman.github.io/verikun/reference/auto-wait/) · [Global flags](https://ddikman.github.io/verikun/reference/global-flags/) · [Exit codes](https://ddikman.github.io/verikun/reference/exit-codes/) · [Environment variables](https://ddikman.github.io/verikun/reference/environment-variables/) · [Reports & test runs](https://ddikman.github.io/verikun/reference/reports-and-test-runs/) · [Device state](https://ddikman.github.io/verikun/reference/device-state/) · [Device claims](https://ddikman.github.io/verikun/reference/device-claims/) · [Screenshots](https://ddikman.github.io/verikun/reference/screenshots/) · [AI plans & models](https://ddikman.github.io/verikun/reference/ai-plans/) |
66
+ | **Reference** | [Commands](https://ddikman.github.io/verikun/reference/commands/) · [Selectors](https://ddikman.github.io/verikun/reference/selectors/) · [Auto-wait](https://ddikman.github.io/verikun/reference/auto-wait/) · [Global flags](https://ddikman.github.io/verikun/reference/commands/#global-flags) · [Exit codes](https://ddikman.github.io/verikun/reference/exit-codes/) · [Environment variables](https://ddikman.github.io/verikun/reference/environment-variables/) · [Reports & test runs](https://ddikman.github.io/verikun/reference/reports-and-test-runs/) · [Device state](https://ddikman.github.io/verikun/reference/device-state/) · [Device claims](https://ddikman.github.io/verikun/reference/device-claims/) · [Screenshots](https://ddikman.github.io/verikun/reference/screenshots/) · [AI plans & models](https://ddikman.github.io/verikun/reference/ai-plans/) |
67
67
  | **Internals** | [Architecture](https://ddikman.github.io/verikun/internals/architecture/) · [Core principles](https://ddikman.github.io/verikun/internals/core-principles/) · [Plan IR & the replay engine](https://ddikman.github.io/verikun/internals/plan-ir-and-engine/) · [Contracts](https://ddikman.github.io/verikun/internals/contracts/) · [Contributing](https://ddikman.github.io/verikun/internals/contributing/) |
68
68
 
69
69
  **Exit codes**, since they are the machine contract everything else rests on: `0` success · `1` not found / assertion failed / timeout · `2` usage error, ambiguous selector, or a device another job is driving · `3` environment error. Data goes to stdout; diagnostics to stderr. [Full contract](https://ddikman.github.io/verikun/reference/exit-codes/). Parallel agents share a host-level [device claim](https://ddikman.github.io/verikun/reference/device-claims/) so two jobs do not silently land on the same phone.
@@ -83,17 +83,7 @@ npm test # type-check + the unit suite (no device needed)
83
83
  npm link # optional: put `verikun` and `vk` on your PATH
84
84
  ```
85
85
 
86
- Zero runtime dependencies; the only dev dependencies are `typescript` and `@types/node`. The full contributor guide — the test loop, the Flutter device fixture, versioning and releasing — is in [Contributing](https://ddikman.github.io/verikun/internals/contributing/).
87
-
88
- To work on the [documentation site](https://ddikman.github.io/verikun/), which is an Astro Starlight project in `docs/` with its own dependencies (so `npm install` above does not cover it):
89
-
90
- ```sh
91
- npm run docs:install # once — installs the site's dependencies (needs Node >= 22.12)
92
- npm run docs # dev server with live reload at http://localhost:4321/verikun/
93
- npm run docs:build # production build; fails on a dead internal link
94
- ```
95
-
96
- Open **http://localhost:4321/verikun/** — the site is served under a base path, so bare `localhost:4321` is a 404. Astro needs **Node ≥ 22.12** while the CLI supports Node ≥ 18, so `nvm use 22` first if your shell's Node is older.
86
+ Zero runtime dependencies; the only dev dependencies are `typescript` and `@types/node`. The full contributor guide — the test loop, the Flutter device fixture, versioning and releasing, and running the documentation site locally — is [Contributing](https://ddikman.github.io/verikun/internals/contributing/).
97
87
 
98
88
  ## License
99
89
 
@@ -2,7 +2,7 @@
2
2
  Object.defineProperty(exports, "__esModule", { value: true });
3
3
  exports.ClaudeProvider = void 0;
4
4
  const errors_1 = require("../errors");
5
- const format_1 = require("../ui/format");
5
+ const wait_1 = require("../wait");
6
6
  const ir_1 = require("./ir");
7
7
  const provider_1 = require("./provider");
8
8
  const grammar_1 = require("./grammar");
@@ -18,8 +18,6 @@ const DEFAULT_REQUEST_TIMEOUT_MS = 120_000;
18
18
  // effort (output_config.effort) is rejected by Haiku 4.5; only send it for models
19
19
  // that accept it.
20
20
  const EFFORT_MODELS = new Set(['claude-opus-4-8', 'claude-sonnet-4-6', 'claude-fable-5']);
21
- const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
22
- const backoffMs = (attempt) => Math.min(1000 * 2 ** (attempt - 1), 15000);
23
21
  class ClaudeProvider {
24
22
  opts;
25
23
  constructor(opts) {
@@ -30,27 +28,8 @@ class ClaudeProvider {
30
28
  return { plan: (0, ir_1.parsePlan)(json), usage };
31
29
  }
32
30
  async repair(ctx) {
33
- const parts = [
34
- 'FAILED STEP: ' + JSON.stringify(ctx.failedStep),
35
- 'FAILURE: ' + ctx.reason,
36
- ];
37
- if (ctx.candidates && ctx.candidates.length) {
38
- parts.push(`The selector matched ${ctx.candidates.length} elements (ambiguous) — pick a more specific selector for the SAME intended element, or give_up if none of them is it.`);
39
- }
40
- parts.push('CURRENT SCREEN:\n' + (0, format_1.formatCompact)(ctx.hierarchy));
41
- const { json, usage } = await this.call(grammar_1.REPAIR_GRAMMAR, parts.join('\n\n'), ir_1.REPAIR_DECISION_JSON_SCHEMA, 1024);
42
- const decision = (json ?? {});
43
- if (decision.decision === 'give_up') {
44
- return {
45
- replaceStep: null,
46
- declineReason: decision.reason?.trim() || 'no element on the current screen matches the step intent',
47
- usage,
48
- };
49
- }
50
- // Hand the proposed leaf back UNVALIDATED — engine.ts validates every repair against
51
- // the grammar before splicing (it is the execution trust boundary and can't assume a
52
- // provider validated). A missing/invalid step is rejected there as a failed repair.
53
- return { replaceStep: (decision.step ?? null), usage };
31
+ const { json, usage } = await this.call(grammar_1.REPAIR_GRAMMAR, (0, provider_1.repairUserPrompt)(ctx), ir_1.REPAIR_DECISION_JSON_SCHEMA, 1024);
32
+ return (0, provider_1.repairDecision)(json, usage);
54
33
  }
55
34
  async call(system, user, schema, maxTokens) {
56
35
  const outputConfig = { format: { type: 'json_schema', schema } };
@@ -111,7 +90,7 @@ class ClaudeProvider {
111
90
  catch (e) {
112
91
  if (attempt++ >= maxRetries)
113
92
  throw new errors_1.CliError(`Anthropic API request failed: ${e.message}`, 3);
114
- await sleep(backoffMs(attempt));
93
+ await (0, wait_1.sleep)((0, wait_1.backoffMs)(attempt));
115
94
  continue;
116
95
  }
117
96
  finally {
@@ -122,7 +101,7 @@ class ClaudeProvider {
122
101
  // Retry 429 + 5xx with backoff, honoring Retry-After (no SDK to do it for us).
123
102
  if ((res.status === 429 || res.status >= 500) && attempt++ < maxRetries) {
124
103
  const retryAfter = Number(res.headers.get('retry-after'));
125
- await sleep(Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter * 1000 : backoffMs(attempt));
104
+ await (0, wait_1.sleep)(Number.isFinite(retryAfter) && retryAfter > 0 ? retryAfter * 1000 : (0, wait_1.backoffMs)(attempt));
126
105
  continue;
127
106
  }
128
107
  const errText = await res.text().catch(() => '');
@@ -8,7 +8,6 @@ const node_fs_1 = require("node:fs");
8
8
  const node_path_1 = require("node:path");
9
9
  const node_os_1 = require("node:os");
10
10
  const errors_1 = require("../errors");
11
- const format_1 = require("../ui/format");
12
11
  const ir_1 = require("./ir");
13
12
  const provider_1 = require("./provider");
14
13
  const grammar_1 = require("./grammar");
@@ -153,23 +152,8 @@ class CliProvider {
153
152
  return { plan: (0, ir_1.parsePlan)(json), usage: {} };
154
153
  }
155
154
  async repair(ctx) {
156
- const parts = ['FAILED STEP: ' + JSON.stringify(ctx.failedStep), 'FAILURE: ' + ctx.reason];
157
- if (ctx.candidates && ctx.candidates.length) {
158
- parts.push(`The selector matched ${ctx.candidates.length} elements (ambiguous) — pick a more specific selector for the SAME intended element, or give_up if none of them is it.`);
159
- }
160
- parts.push('CURRENT SCREEN:\n' + (0, format_1.formatCompact)(ctx.hierarchy));
161
- const json = this.call(grammar_1.REPAIR_GRAMMAR, parts.join('\n\n'), ir_1.REPAIR_DECISION_JSON_SCHEMA);
162
- const decision = (json ?? {});
163
- if (decision.decision === 'give_up') {
164
- return {
165
- replaceStep: null,
166
- declineReason: decision.reason?.trim() || 'no element on the current screen matches the step intent',
167
- usage: {},
168
- };
169
- }
170
- // Hand the proposed leaf back UNVALIDATED — engine.ts validates every repair against the
171
- // grammar before splicing (it is the execution trust boundary), exactly like the API providers.
172
- return { replaceStep: (decision.step ?? null), usage: {} };
155
+ // usage:{} for the same reason as compile — billed to the subscription, not per token.
156
+ return (0, provider_1.repairDecision)(this.call(grammar_1.REPAIR_GRAMMAR, (0, provider_1.repairUserPrompt)(ctx), ir_1.REPAIR_DECISION_JSON_SCHEMA), {});
173
157
  }
174
158
  /** Spawn the CLI once and return the parsed JSON object it produced. Synchronous (spawnSync);
175
159
  * the async method wrappers satisfy the Promise-returning AgentProvider seam. */
@@ -6,6 +6,7 @@ const node_crypto_1 = require("node:crypto");
6
6
  const selector_1 = require("../ui/selector");
7
7
  const state_support_1 = require("../ui/state-support");
8
8
  const errors_1 = require("../errors");
9
+ const wait_1 = require("../wait");
9
10
  const ir_1 = require("./ir");
10
11
  /** An outcome is environment-flavoured if it carries an exit-3 CliError, or simply
11
12
  * reported code 3 — the latter also catches a remote step whose error crossed the
@@ -70,11 +71,10 @@ function isHealable(outcome) {
70
71
  }
71
72
  /** Default wall-clock ceiling for a whole `vk ai` run (overridable via --timeout). */
72
73
  exports.DEFAULT_RUN_TIMEOUT_MS = 15 * 60 * 1000;
73
- const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
74
74
  /** How long a CONDITIONAL guard (`if-present`) waits for its selector to show up
75
75
  * before concluding "absent". Interstitials animate in: a permission dialog or promo
76
76
  * panel typically lands a few hundred ms after the transition that triggers it. Every
77
- * selector-resolving leaf command already auto-waits ~5s (cli.ts resolveOneWaiting),
77
+ * selector-resolving leaf command already auto-waits ~5s (commands/auto-wait.ts),
78
78
  * so before this window existed a guard was strictly LESS patient than a bare `tap` —
79
79
  * and an optional dialog could be missed by the very construct meant to catch it.
80
80
  * Kept far below the leaf 5s because an ABSENT guard pays this window, every time, and
@@ -88,6 +88,32 @@ const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
88
88
  exports.DEFAULT_GUARD_SETTLE_MS = 1500;
89
89
  /** Re-dump cadence inside a guard's settle window. */
90
90
  const GUARD_POLL_MS = 150;
91
+ /**
92
+ * How long a guard keeps looking at a screen it cannot read at all, when the reason is
93
+ * NoWindowError — the app was force-stopped or is mid-launch and has genuinely not drawn.
94
+ *
95
+ * Separate from `settleMs`, which answers "how long before I believe this selector is
96
+ * absent?". This answers "how long before I believe there is no screen to ask?" — a
97
+ * different question, and the only one whose wrong answer aborts the whole run.
98
+ *
99
+ * MEASURED (issue #80), post-`vk launch`, time until the first hierarchy read succeeds:
100
+ * emulator-5554 (stock reads ~2s): 2027-2379ms over 10 launches
101
+ * SM-A415F, physical (companion ~0.3s): 4746-6219ms over 15 launches
102
+ * So the "clears within a second or two" in NoWindowError's own doc is optimistic by 2.5x on
103
+ * real hardware; the emulator only looks compliant because one slow read already outlasts the
104
+ * gap. 10s clears the worst observed by ~1.6x. Below that the constant sits INSIDE the
105
+ * measured distribution, which is the one place it must not be.
106
+ *
107
+ * Deliberately MORE patient than a leaf command's 5s auto-wait, which is the opposite of
108
+ * DEFAULT_GUARD_SETTLE_MS's reasoning, because the asymmetry is real: a leaf that needs
109
+ * longer takes `--wait`, whereas a guard's patience is internal and a test author cannot
110
+ * reach it. Being generous costs at most this long ONCE per present() call, on a run that is
111
+ * already failing — against a 15-minute default run timeout, and clamped by it.
112
+ *
113
+ * Not configurable on purpose: a dial here is one more thing to explain, and every value a
114
+ * user might pick is worse than the measurement.
115
+ */
116
+ const NO_WINDOW_GRACE_MS = 10_000;
91
117
  /** Consecutive identical screen snapshots before a loop is believed to be stuck.
92
118
  *
93
119
  * This check is a TIME SAVER and nothing more. A loop already fails when its exit
@@ -179,6 +205,11 @@ async function runPlan(plan, deps) {
179
205
  * probe a loop-exit check needs.
180
206
  * So one dump attempt always happens regardless of the window.
181
207
  *
208
+ * A THIRD clock sits beside both: a screen that cannot be read because the app has not
209
+ * drawn (NoWindowError) is retried against NO_WINDOW_GRACE_MS, not against `settleMs`.
210
+ * That is deliberately independent — "is this selector absent?" and "is there a screen to
211
+ * ask at all?" are different questions, and only the second one aborts the run.
212
+ *
182
213
  * Throws GuardBlindError when the window closes having NEVER once read the screen
183
214
  * and the failure was an environment error — see that class for why. */
184
215
  const present = async (selector, settleMs) => {
@@ -198,6 +229,10 @@ async function runPlan(plan, deps) {
198
229
  if (deps.platform)
199
230
  (0, state_support_1.assertStateSupported)(sel, deps.platform);
200
231
  const deadline = Date.now() + Math.max(0, settleMs);
232
+ // A screen that cannot be read AT ALL gets its own, longer clock — see NO_WINDOW_GRACE_MS.
233
+ // Clamped by the run deadline so a guard can never push a run past --timeout: the grace
234
+ // exists to spend budget the caller already has, never to invent more.
235
+ const noWindowDeadline = Math.min(Date.now() + NO_WINDOW_GRACE_MS, deps.deadline ?? Infinity);
201
236
  // A non-zero window must buy at least one SECOND look, independent of the clock.
202
237
  // Measured on emulator-5554: one uiautomator dump costs ~2.4s, which already exceeds
203
238
  // a 1.5s window — so a purely time-boxed loop returns after a single dump and the
@@ -235,14 +270,31 @@ async function runPlan(plan, deps) {
235
270
  return true;
236
271
  const remaining = deadline - Date.now();
237
272
  if (looks >= minLooks && remaining <= 0) {
273
+ // Nothing has been readable yet, and the reason is that the app has not drawn. That is
274
+ // an observation about the SCREEN, not a broken machine, and it clears on its own — so
275
+ // keep looking on the no-window clock instead of killing the run. MEASURED (#80): the
276
+ // two attempts above span ~75ms against a gap of 4.7-6.2s on a physical device, so
277
+ // without this a `repeat` with minutes of budget gives up in under a tenth of a second.
278
+ //
279
+ // Gated on `everRead` so this can only ever extend patience for an UNREADABLE screen.
280
+ // One successful read — even an empty tree — and the ordinary semantics resume exactly:
281
+ // settleMs=0 is still a single-shot probe. It must never make a merely ABSENT selector
282
+ // more patient, or every guard silently costs 10s.
283
+ if (!everRead && lastErr instanceof errors_1.NoWindowError && Date.now() < noWindowDeadline) {
284
+ await (0, wait_1.sleep)(GUARD_POLL_MS);
285
+ continue;
286
+ }
238
287
  // The window closed having NEVER once read the screen, because the environment is
239
288
  // broken. Answering "absent" here is a lie that silently skips the body — and a
240
289
  // guard-heavy plan would then finish fully GREEN having executed nothing.
290
+ //
291
+ // A no-window that outlives its grace lands here too, and still aborts: at that point
292
+ // the app really is gone, and reporting "absent" would be the same false green.
241
293
  if (!everRead && (0, errors_1.isEnvError)(lastErr))
242
294
  throw new GuardBlindError(selector, lastErr);
243
295
  return false;
244
296
  }
245
- await sleep(Math.max(0, Math.min(GUARD_POLL_MS, remaining)));
297
+ await (0, wait_1.sleep)(Math.max(0, Math.min(GUARD_POLL_MS, remaining)));
246
298
  }
247
299
  };
248
300
  const runLeaf = (leaf) => {
@@ -397,7 +449,7 @@ async function runPlan(plan, deps) {
397
449
  // to happen. Fail where the information is.
398
450
  return { status: 'fail', where, reason: `read found no element matching '${selector}' (nothing to store in ctx.${node.into})` };
399
451
  }
400
- await sleep(Math.max(0, Math.min(GUARD_POLL_MS, remaining)));
452
+ await (0, wait_1.sleep)(Math.max(0, Math.min(GUARD_POLL_MS, remaining)));
401
453
  }
402
454
  }
403
455
  async function walkNode(node, where, replace, guard) {
package/dist/agent/ir.js CHANGED
@@ -63,6 +63,30 @@ exports.KNOWN_COMMANDS = new Set([
63
63
  // non-recordable selector-resolver here (e.g. `find`) would corrupt an unrelated step.
64
64
  ]);
65
65
  exports.DEFAULT_LOOP_CAP = 25;
66
+ // The two SIMPLE control nodes appear at every unrolled level (see innerControls), so
67
+ // their shape is written once. Key order is the wire format — keep it.
68
+ const ifPresentSchema = (body) => ({
69
+ type: 'object',
70
+ additionalProperties: false,
71
+ required: ['type', 'selector', 'body'],
72
+ properties: {
73
+ type: { type: 'string', enum: ['if-present'] },
74
+ selector: { type: 'string' },
75
+ body,
76
+ },
77
+ });
78
+ const whilePresentSchema = (body) => ({
79
+ type: 'object',
80
+ additionalProperties: false,
81
+ required: ['type', 'selector', 'cap', 'body'],
82
+ properties: {
83
+ type: { type: 'string', enum: ['while-present'] },
84
+ selector: { type: 'string' },
85
+ bind: { type: 'string' },
86
+ cap: { type: 'integer' },
87
+ body,
88
+ },
89
+ });
66
90
  /**
67
91
  * JSON Schema for `output_config.format` so the model returns a guaranteed-valid
68
92
  * Plan. Deliberately NON-RECURSIVE: instead of a `Plan -> node -> Plan` cycle
@@ -73,16 +97,7 @@ const stepItems = (bodyItems) => ({
73
97
  anyOf: [
74
98
  leafSchema(),
75
99
  readSchema(),
76
- {
77
- type: 'object',
78
- additionalProperties: false,
79
- required: ['type', 'selector', 'body'],
80
- properties: {
81
- type: { type: 'string', enum: ['if-present'] },
82
- selector: { type: 'string' },
83
- body: { type: 'array', items: bodyItems },
84
- },
85
- },
100
+ ifPresentSchema({ type: 'array', items: bodyItems }),
86
101
  {
87
102
  type: 'object',
88
103
  additionalProperties: false,
@@ -112,18 +127,7 @@ const stepItems = (bodyItems) => ({
112
127
  body: { type: 'array', items: bodyItems },
113
128
  },
114
129
  },
115
- {
116
- type: 'object',
117
- additionalProperties: false,
118
- required: ['type', 'selector', 'cap', 'body'],
119
- properties: {
120
- type: { type: 'string', enum: ['while-present'] },
121
- selector: { type: 'string' },
122
- bind: { type: 'string' },
123
- cap: { type: 'integer' },
124
- body: { type: 'array', items: bodyItems },
125
- },
126
- },
130
+ whilePresentSchema({ type: 'array', items: bodyItems }),
127
131
  ],
128
132
  });
129
133
  /** The innermost level: leaves, reads, and the two SIMPLE control nodes — `if-present`
@@ -148,30 +152,7 @@ const stepItems = (bodyItems) => ({
148
152
  * here is what would actually cost the schema. */
149
153
  const innerControls = () => {
150
154
  const leafBody = { type: 'array', items: { anyOf: [leafSchema(), readSchema()] } };
151
- return [
152
- {
153
- type: 'object',
154
- additionalProperties: false,
155
- required: ['type', 'selector', 'body'],
156
- properties: {
157
- type: { type: 'string', enum: ['if-present'] },
158
- selector: { type: 'string' },
159
- body: leafBody,
160
- },
161
- },
162
- {
163
- type: 'object',
164
- additionalProperties: false,
165
- required: ['type', 'selector', 'cap', 'body'],
166
- properties: {
167
- type: { type: 'string', enum: ['while-present'] },
168
- selector: { type: 'string' },
169
- bind: { type: 'string' },
170
- cap: { type: 'integer' },
171
- body: leafBody,
172
- },
173
- },
174
- ];
155
+ return [ifPresentSchema(leafBody), whilePresentSchema(leafBody)];
175
156
  };
176
157
  const LEAF_ONLY_ITEMS = { anyOf: [leafSchema(), readSchema(), ...innerControls()] };
177
158
  exports.PLAN_JSON_SCHEMA = {