verikun 0.29.0 → 0.29.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,14 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.29.1] - 2026-09-26
10
+
11
+ ### Fixed
12
+ - **`vk ai`** no longer guesses `id:` selectors; an element the test names only by label compiles
13
+ to `text:` — write the id into the test to pin it. ([#148])
14
+
15
+ [#148]: https://github.com/ddikman/verikun/issues/148
16
+
9
17
  ## [0.29.0] - 2026-09-16
10
18
 
11
19
  ### Changed
@@ -118,7 +118,7 @@ Do NOT hard-code a run of indices you were not told the length of. If the prose
118
118
  varies per run — emit a while-present over {{ctx.i}} rather than tap _0, _1, _2, _3.
119
119
  A hard-coded list is right only when the prose states the exact count.
120
120
 
121
- SELECTORS (the engine auto-heals case/whitespace/partial, so prefer stable identifiers):
121
+ SELECTORS (the engine auto-heals case/whitespace/partial, so prefer an id the test gives you):
122
122
  @login resource-id 'login' (shorthand for id:login)
123
123
  id:login resource-id (full, suffix, or short)
124
124
  text:Sign in visible text (case-insensitive)
@@ -156,7 +156,17 @@ RULES:
156
156
  below the fold — that is now redundant. Emit an explicit swipe only when the SCROLLING
157
157
  ITSELF is what the test asks for ("scroll the feed three times"), or to reveal content
158
158
  that is not in the hierarchy until it is built (an infinite/lazy list).
159
- - Prefer resource-id / accessibility selectors over visible text where possible.
159
+ - IDS COME FROM THE TEST. Use @x / id:x only with an id the test writes (or one a PRIOR PLAN
160
+ shown to you already uses). You cannot see the app, so never compose an id from a label, a
161
+ description, or the pattern of the app's other ids: a made-up id matches nothing, and a
162
+ wait/assert on it FAILS the test with no repair. An element the test names only by what it
163
+ says ("tap Sign in", "the Home tab") is text:<those words> — forgiving, and it also matches an
164
+ accessibility label.
165
+ - An outcome the test states ("the tab bar (Home, Search, Profile) appears") is checked on what
166
+ it NAMES — the id it gives, or the words it quotes or lists — never on an id you made up. A
167
+ description that names nothing you could select without inventing it ("it shows the
168
+ signed-in user's name") is context, not a step; so is rationale (why a step exists, what a
169
+ field already holds).
160
170
  - Translate the test literally and minimally: do not invent ACTION steps (tap/text/swipe/key/assert)
161
171
  the prose does not imply. The ONE exception is screenshot — insert screenshot steps liberally as
162
172
  post-run review evidence: after each screen transition (launch, a navigation tap, a submit, a
@@ -2,7 +2,7 @@
2
2
  // Compile-fidelity lint: does the plan the model produced still say what the prose said?
3
3
  //
4
4
  // The model is a compiler, and this is the compiler's own sanity check. It exists because
5
- // compilation is NONDETERMINISTIC in a way that is invisible until a run fails, and three
5
+ // compilation is NONDETERMINISTIC in a way that is invisible until a run fails, and four
6
6
  // failure modes showed up repeatedly against a real suite:
7
7
  //
8
8
  // 1. An explicit directive silently vanishes. The same prose ("Launch the app WITH ITS
@@ -18,8 +18,12 @@
18
18
  // against later builds. A test exercising none of its subject reporting success is the
19
19
  // worst failure mode a testing tool has, which is why the two rules that detect it are
20
20
  // the only FATAL findings here.
21
+ // 4. A selector names an id the test never gave (issue #148). The compiler cannot see the
22
+ // app, so it composed one from the app's naming pattern, and 3 of 4 such guesses named
23
+ // nothing in the app. A wait/assert is never repaired, so the test went red — and each
24
+ // recompile guessed differently, so it flaked.
21
25
  //
22
- // 1 and 2 are cheap to detect and cheap to fix: hand the finding back to the model and let
26
+ // 1, 2 and 4 are cheap to detect and cheap to fix: hand the finding back to the model and let
23
27
  // it compile once more. That is far better than the alternative, which is a plan that is
24
28
  // quietly wrong and burns a device run to say so. 3 gets the same guided recompile, but a
25
29
  // plan that STILL does not cover its test must never run — see `fatal` below.
@@ -34,9 +38,11 @@ exports.instructionLines = instructionLines;
34
38
  exports.instructionUnits = instructionUnits;
35
39
  exports.actionNodes = actionNodes;
36
40
  exports.tailAnchors = tailAnchors;
41
+ exports.ungroundedIds = ungroundedIds;
37
42
  exports.lintPlan = lintPlan;
38
43
  exports.looksTruncated = looksTruncated;
39
44
  const ir_1 = require("./ir");
45
+ const selector_1 = require("../ui/selector");
40
46
  /** Walk every node in the plan, including control-node bodies. */
41
47
  function* walk(nodes) {
42
48
  for (const n of nodes) {
@@ -257,15 +263,95 @@ function planMentions(tokens, anchor) {
257
263
  return b.length >= 3 && (b.includes(a) || a.includes(b));
258
264
  });
259
265
  }
266
+ // --- ids: does every id the plan selects by come from the prose? (issue #148) ------------
267
+ /** Commands whose FIRST positional is a selector. The rest of `text`'s positionals, and every
268
+ * positional of `type`/`launch`/…, is data — `type @handle` types a handle, it selects nothing. */
269
+ const SELECTOR_FIRST = new Set(['tap', 'click', 'text', 'wait', 'assert']);
270
+ /** Every string the engine resolves as a selector — unlike `planTokens`, which also returns data. */
271
+ function selectorsOf(plan) {
272
+ const out = [];
273
+ for (const n of walk(plan.steps)) {
274
+ if (n.type === 'command') {
275
+ if (SELECTOR_FIRST.has(n.command) && n.positionals[0])
276
+ out.push(n.positionals[0]);
277
+ if (n.command === 'swipe' || n.command === 'scroll') {
278
+ for (const f of n.flags)
279
+ if (f.name === 'on')
280
+ out.push(f.value);
281
+ }
282
+ }
283
+ else if (n.type === 'when') {
284
+ out.push(...n.branches.map((b) => b.selector));
285
+ }
286
+ else {
287
+ out.push(n.selector);
288
+ }
289
+ }
290
+ return out;
291
+ }
292
+ /** The id values (`@x` / `id:x`) the plan selects by, parsed by the engine's own parser so the
293
+ * two can never disagree about what an id selector is — trailing state modifiers included. */
294
+ function planIds(plan) {
295
+ const out = [];
296
+ for (const raw of selectorsOf(plan)) {
297
+ let sel;
298
+ try {
299
+ sel = (0, selector_1.parseSelector)(raw);
300
+ }
301
+ catch {
302
+ continue; // an empty value or contradictory modifiers: a runtime error, not a guess
303
+ }
304
+ if (sel.kind === 'id' && !out.includes(sel.value))
305
+ out.push(sel.value);
306
+ }
307
+ return out;
308
+ }
309
+ /** `com.app:id/login` and `login` name the same element. */
310
+ const RESOURCE_PKG_RE = /^[\w.]+:id\//;
311
+ /** `{{ctx.i}}` / `{{env.X}}` are filled at run time, so they are never part of a guess. */
312
+ const PLACEHOLDER_RE = /\{\{[^}]*\}\}/;
313
+ /** The literal parts of an id worth judging. Under three characters a substring matches
314
+ * anything, so a part that short has no opinion. */
315
+ const idParts = (id) => id
316
+ .replace(RESOURCE_PKG_RE, '')
317
+ .split(PLACEHOLDER_RE)
318
+ .map(normToken)
319
+ .filter((p) => p.length >= 3);
320
+ /** Every word of the prose, and every quoted span whole — an iOS identifier may hold spaces. */
321
+ function proseTokens(nl) {
322
+ const out = [];
323
+ for (const m of nl.matchAll(ANCHOR_RE))
324
+ out.push(m[1] ?? m[2] ?? m[3] ?? '');
325
+ out.push(...nl.split(/[\s`"“”'‘’()[\]{}<>,;!?]+/));
326
+ return out.map(normToken).filter((t) => t.length > 0);
327
+ }
328
+ /**
329
+ * The ids the plan selects by that neither the prose nor `seed` gives — each one a guess.
330
+ *
331
+ * LENIENT on purpose, since a false positive costs a recompile: an id is grounded when every
332
+ * literal part of it (package qualifier and placeholders dropped, case and punctuation ignored)
333
+ * sits inside ONE token of the prose, so `@spinner` is grounded by `vk_spinner` and
334
+ * `id:option_{{ctx.i}}` by `option_0`. One token, never two: `settings_tab` built from the words
335
+ * "Settings tab" is precisely how a guess is made.
336
+ *
337
+ * `seed` is a prior plan the compiler was told to reuse. A green run re-persists the HEALED plan
338
+ * and a repair takes its id off a live screen, so an id a seed already uses is grounded too —
339
+ * pass one only where it can hold a repair (see `compileFromSegments` in ../cli.ts).
340
+ */
341
+ function ungroundedIds(nl, plan, seed) {
342
+ const ground = [...proseTokens(nl), ...(seed ? planIds(seed).flatMap(idParts) : [])];
343
+ return planIds(plan).filter((id) => idParts(id).some((part) => !ground.some((t) => t.includes(part))));
344
+ }
260
345
  /**
261
346
  * Check a compiled plan against the prose it came from.
262
347
  *
263
348
  * @param nl the natural-language test, verbatim
264
349
  * @param plan the plan the model just produced
350
+ * @param seed the prior plan the model was handed to adapt, if any — its ids are not guesses
265
351
  * @returns findings; empty means the plan is consistent with the prose. A finding with
266
352
  * `fatal` set means the plan does not cover the test and must not be run.
267
353
  */
268
- function lintPlan(nl, plan) {
354
+ function lintPlan(nl, plan, seed) {
269
355
  const findings = [];
270
356
  if (FRESH_START_RE.test(nl) && !hasLeafWithFlag(plan, 'launch', 'clear')) {
271
357
  findings.push({
@@ -281,6 +367,21 @@ function lintPlan(nl, plan) {
281
367
  '(skip when absent), or behind when (when the screen is one of several known kinds).',
282
368
  });
283
369
  }
370
+ // Not fatal: a guess that happens to be right still runs, and a wrong one fails its step
371
+ // loudly, so the guided recompile is the remedy. The exception is an ABSENCE check (`--gone`,
372
+ // a guard that skips): it passes on any selector that never matches, a wrong label as much as
373
+ // a guessed id, so making this rule fatal would not close it.
374
+ const guessed = ungroundedIds(nl, plan, seed);
375
+ if (guessed.length > 0) {
376
+ const named = guessed.slice(0, 5).map((id) => JSON.stringify(id)).join(', ');
377
+ findings.push({
378
+ message: `The plan selects by id ${named}${guessed.length > 5 ? ` and ${guessed.length - 5} more` : ''}, but the test ` +
379
+ `never gives ${guessed.length === 1 ? 'that id' : 'those ids'}. You cannot see the app, so an id the test does not state is a guess — it matches ` +
380
+ 'nothing, and a wait/assert on it fails the test with no repair. Use an id only when the test writes it; ' +
381
+ 'select an element the test names by its label with text:<label>, and do not check something the test ' +
382
+ 'names nothing selectable for.',
383
+ });
384
+ }
284
385
  if (!coverageChecksEnabled())
285
386
  return findings;
286
387
  const units = instructionUnits(nl);
package/dist/cli.js CHANGED
@@ -1905,7 +1905,14 @@ async function compileFromSegments(segments, key, opts, cost, provider) {
1905
1905
  // FREE, and refusing it over a ceiling we were never about to spend against would fail
1906
1906
  // a test for somebody else's tokens.
1907
1907
  assertBudgetForCompile(cost, opts.maxCostUsd, `compiling ${where} — the test is only partly compiled`);
1908
- const seed = seedPlan(segKey, where);
1908
+ // A section's entry is only ever raw compile output — a green run re-persists the WHOLE
1909
+ // test's key, never a section's — so an id in it that its own prose never gives can only be
1910
+ // a guess (issue #148), and handing it back as "reuse this" would keep it alive.
1911
+ let seed = seedPlan(segKey, where);
1912
+ if (seed && (0, lint_1.ungroundedIds)(seg.text, seed.plan).length > 0) {
1913
+ (0, output_1.err)(`[ai] ${where}: ignoring a prior plan that guesses ids its prose never gives — compiling fresh`);
1914
+ seed = null;
1915
+ }
1909
1916
  (0, output_1.err)(`[ai] ${where}: compiling with ${opts.model}…${lock.degraded ? ` (${lock.degraded})` : ''}`);
1910
1917
  let compiled;
1911
1918
  try {
@@ -1939,6 +1946,14 @@ async function compileFromSegments(segments, key, opts, cost, provider) {
1939
1946
  (0, output_1.err)(`[ai] ${where}: the compiled section does not cover its prose (${compiled.plan.steps.length} step(s)) — not caching it; compiling the test as one instead`);
1940
1947
  return null; // the finally below releases the lock
1941
1948
  }
1949
+ // A guessed id in a FRAGMENT is issue #148 spliced into every test that includes it. Same
1950
+ // answer as a short section: keep it out of the cache and compile the test whole, where the
1951
+ // lint hands the guess back for one guided recompile.
1952
+ const guessed = (0, lint_1.ungroundedIds)(seg.text, compiled.plan);
1953
+ if (guessed.length > 0) {
1954
+ (0, output_1.err)(`[ai] ${where}: the compiled section selects by ids its prose never gives (${guessed.join(', ')}) — not caching it; compiling the test as one instead`);
1955
+ return null; // the finally below releases the lock
1956
+ }
1942
1957
  // INSIDE the lock and before the release: a waiter re-reads the cache the instant the
1943
1958
  // lock disappears, so releasing first would hand it a miss and buy the second compile
1944
1959
  // this whole mechanism exists to prevent.
@@ -2013,7 +2028,10 @@ async function obtainPlan(key, file, opts, cost, provider, segments = []) {
2013
2028
  // sometimes only the test's opening compiles at all. One guided retry is much cheaper
2014
2029
  // than discovering it as a device-run failure several steps later, or (worse, for a
2015
2030
  // truncation) as a pass. Budget is re-checked HERE: the first attempt is already billed.
2016
- let remaining = (0, lint_1.lintPlan)(key.nl, compiled.plan);
2031
+ //
2032
+ // The seed goes in too: a green run re-persists the HEALED plan, whose repaired ids came off a
2033
+ // live screen, so flagging them would push every new build off a selector the device confirmed.
2034
+ let remaining = (0, lint_1.lintPlan)(key.nl, compiled.plan, seed?.plan);
2017
2035
  if (remaining.length > 0) {
2018
2036
  const feedback = remaining.map((f) => `- ${f.message}`).join('\n');
2019
2037
  (0, output_1.err)(`[ai] compiled plan does not match the test — recompiling once:\n${feedback}`);
@@ -2029,7 +2047,7 @@ async function obtainPlan(key, file, opts, cost, provider, segments = []) {
2029
2047
  retryFeedback: feedback,
2030
2048
  });
2031
2049
  cost.add(retry.usage, 'compile');
2032
- const still = (0, lint_1.lintPlan)(key.nl, retry.plan);
2050
+ const still = (0, lint_1.lintPlan)(key.nl, retry.plan, seed?.plan);
2033
2051
  // Keep the retry either way: it was compiled with strictly more information. If it
2034
2052
  // still trips the lint, say so rather than pretending the plan is clean — and for a
2035
2053
  // coverage finding, do not claim it will run, because the gate below rejects it.
package/dist/version.js CHANGED
@@ -3,4 +3,4 @@ Object.defineProperty(exports, "__esModule", { value: true });
3
3
  exports.VERSION = void 0;
4
4
  // GENERATED by scripts/gen-version.mjs from package.json's "version" at build time
5
5
  // (the `prebuild` script). Do NOT edit by hand; bump package.json instead.
6
- exports.VERSION = '0.29.0';
6
+ exports.VERSION = '0.29.1';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "verikun",
3
- "version": "0.29.0",
3
+ "version": "0.29.1",
4
4
  "description": "Drive Android emulators/devices and iOS simulators for AI agents: tap, type, swipe, screenshot, and inspect the UI hierarchy by semantic identifiers — like Puppeteer for native apps.",
5
5
  "keywords": [
6
6
  "android",