simframe 0.13.0 → 0.14.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -406,6 +406,7 @@ steer the model is a tool surface the model uses wrong.
406
406
  | `sim_wait` | Waits for the screen to change *and then* settle. |
407
407
  | `sim_look` | **The only tool that returns an image**, capped at 1024 px. For layout, colour, spacing — questions text cannot answer. |
408
408
  | `sim_recall` · `sim_strip` | Look backwards: a text timeline of what happened, or recent frames tiled into one image. |
409
+ | `sim_storage` | **What the app believes**, as opposed to what it drew: its `UserDefaults` and, for React Native, its `AsyncStorage`. Reads the data container off disk, so it answers on a device that is **not running**. |
409
410
  | `sim_capture` · `sim_devices` | Manage capture loops; list simulators. |
410
411
 
411
412
  ### What the screen looks like as text
@@ -585,6 +586,20 @@ simframe screens # what this device has learned
585
586
  simframe goto invoices # walk there, verifying every step
586
587
  ```
587
588
 
589
+ And when the screen and the behaviour disagree, the question is usually not
590
+ about the screen at all:
591
+
592
+ ```bash
593
+ simframe storage # apps with a data container
594
+ simframe storage com.example.myapp # what that app saved
595
+ ```
596
+
597
+ `sim_ui` says what is drawn; `sim_storage` says what the app believes. It reads
598
+ the data container straight off the host filesystem, which means it works on a
599
+ device that is **shut down** — `simctl` cannot do this at all, on any of its own
600
+ paths, once a device stops running.
601
+
602
+
588
603
  Measured on a four-tab tour, `goto` plans and walks three-step routes with every
589
604
  step verified and no model call. It fails rather than guesses: an unknown
590
605
  destination, a query matching two screens equally, or no path of known edges all
@@ -13,6 +13,10 @@ public enum Motion {
13
13
 
14
14
  /// Below this mean absolute difference, two frames are the same picture.
15
15
  public static let stillThreshold = 0.004
16
+
17
+ /// How many moving cells make an animation rather than sensor noise.
18
+ /// Set from the measurements in `state(history:now:)` below.
19
+ public static let minAnimatingCells = 12
16
20
  /// Frames that must agree before the screen counts as settled.
17
21
  ///
18
22
  /// Paired with a duration, because frame count alone is not a measure of
@@ -170,8 +174,35 @@ public enum Motion {
170
174
  let cellThreshold = 0.06
171
175
  let moving = union.filter { $0 > cellThreshold }.count
172
176
  let fraction = Double(moving) / Double(union.count)
173
- // Small and persistent, rather than a screen changing.
174
- if fraction > 0 && fraction < 0.06 {
177
+ // Small and persistent, rather than a screen changing — and big
178
+ // enough to be something.
179
+ //
180
+ // `fraction > 0` meant one cell of 4,608 counted as an animation,
181
+ // and measured on a device that fires constantly on screens where
182
+ // nothing is happening. A field report put it exactly right: a
183
+ // warning that is usually wrong trains the reader to ignore the one
184
+ // that matters.
185
+ //
186
+ // Measured on this device, 2026-09-14, which is what
187
+ // `minAnimatingCells` is set from:
188
+ //
189
+ // static home screen, nothing moving 1-4 cells, on 54% of frames
190
+ // a blinking text caret **2 cells** (1x2), on 100%
191
+ // the testbed's spinner 169 cells (13x13)
192
+ // Maps launching, median 656 cells (41x16)
193
+ // a real spinner, from the field 1,364 cells (44x31)
194
+ //
195
+ // Caret and noise sit at 1-4; the smallest real animation seen is
196
+ // 169. A 42x gap with nothing in it, so the threshold is not
197
+ // delicate — 12 is three times the worst noise and an order of
198
+ // magnitude below the weakest signal.
199
+ //
200
+ // The caret is the case this is really for, and it is why the
201
+ // question could not be settled by reasoning: a caret must never
202
+ // stop a screen from settling, and before this it did — 72 frames
203
+ // out of 72 with `settled: false` on a screen holding nothing but a
204
+ // text cursor.
205
+ if moving >= Self.minAnimatingCells && fraction < 0.06 {
175
206
  animating = boundingBox(of: union, threshold: cellThreshold)
176
207
  }
177
208
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "simframe",
3
- "version": "0.13.0",
3
+ "version": "0.14.1",
4
4
  "mcpName": "io.github.lvlrSajjad/simframe",
5
5
  "description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
6
6
  "keywords": [
@@ -0,0 +1,82 @@
1
+ #!/usr/bin/env node
2
+ // Run a CI step, and tell a sick simulator apart from a failing check.
3
+ //
4
+ // node scripts/ci-device-guard.mjs <udid> -- <command> [args...]
5
+ //
6
+ // **Why this exists, with the number that justifies it.** Over the last 25 CI
7
+ // runs the `integration` job failed 17 times and passed 5. Eight of the nine
8
+ // most recent failures name a device-state condition in their own output —
9
+ // `NSPOSIXErrorDomain code=60`, `the display produced no frame in 60s`, `the
10
+ // second app never launched` — and one of those runs was a **docs-only commit**
11
+ // that changed a single markdown file. Two runs of byte-identical code failed at
12
+ // two different steps.
13
+ //
14
+ // So the job has been answering two questions at once — *does simframe work* and
15
+ // *did this hosted simulator survive twenty minutes* — and the second dominates.
16
+ // A red build that is usually the device is the cry-wolf failure this repo keeps
17
+ // writing items about, and it cost two days of re-reading logs to learn nothing.
18
+ //
19
+ // This does not paper over failures. It distinguishes them: a named device
20
+ // condition gets the cure this project already ships (`simframe revive`) and one
21
+ // retry, exactly as DEFERRED 126 says to; anything else fails on the spot,
22
+ // untouched. The classification is written to the job summary so the *rate*
23
+ // becomes visible instead of arguable.
24
+ import { spawn } from 'node:child_process';
25
+ import fs from 'node:fs';
26
+
27
+ /** Conditions that are the simulator, not the code. Each seen in a real run. */
28
+ const DEVICE_STATE = [
29
+ [/NSPOSIXErrorDomain.*code=?\s*60|Operation timed out/i, 'simctl stopped answering (NSPOSIXErrorDomain 60)'],
30
+ [/did not produce a frame|produced no frame in \d+s/i, 'the daemon is up and the display renders nothing'],
31
+ [/Timeout waiting for screen surfaces|display surface is not answering|display surface could not be read/i, 'the display surface is wedged'],
32
+ [/no frames buffered|capture is wedged/i, 'capture stopped'],
33
+ [/the second app never launched|could not be dispatched/i, 'an app would not launch'],
34
+ ];
35
+
36
+ const udid = process.argv[2];
37
+ const sep = process.argv.indexOf('--');
38
+ if (!udid || sep < 0) {
39
+ console.error('usage: ci-device-guard.mjs <udid> -- <command> [args...]');
40
+ process.exit(2);
41
+ }
42
+ const cmd = process.argv.slice(sep + 1);
43
+
44
+ function run(argv, { capture = true } = {}) {
45
+ return new Promise((resolve) => {
46
+ const p = spawn(argv[0], argv.slice(1), { stdio: capture ? ['inherit', 'pipe', 'pipe'] : 'inherit' });
47
+ let out = '';
48
+ p.stdout?.on('data', (d) => { out += d; process.stdout.write(d); });
49
+ p.stderr?.on('data', (d) => { out += d; process.stderr.write(d); });
50
+ p.on('close', (code) => resolve({ code, out }));
51
+ });
52
+ }
53
+
54
+ const summary = (line) => {
55
+ const f = process.env.GITHUB_STEP_SUMMARY;
56
+ if (f) { try { fs.appendFileSync(f, `${line}\n`); } catch { /* summaries are a nicety */ } }
57
+ };
58
+
59
+ const deviceCause = (text) => DEVICE_STATE.find(([re]) => re.test(text))?.[1] ?? null;
60
+
61
+ const first = await run(cmd);
62
+ if (first.code === 0) process.exit(0);
63
+
64
+ const cause = deviceCause(first.out);
65
+ if (!cause) {
66
+ console.error(`\n (this step failed on its merits, not on the device — not retrying)`);
67
+ summary(`- \`${cmd.join(' ')}\` — **check failed** (exit ${first.code})`);
68
+ process.exit(first.code ?? 1);
69
+ }
70
+
71
+ console.error(`\n (${cause} — DEFERRED 126. Reviving once and running again.)`);
72
+ await run(['node', 'src/cli.js', 'revive', `--device=${udid}`], { capture: false });
73
+ const second = await run(cmd);
74
+ if (second.code === 0) {
75
+ summary(`- \`${cmd.join(' ')}\` — passed after one revive (${cause})`);
76
+ process.exit(0);
77
+ }
78
+ // Twice in a row, on a condition we know the cure for. Reported as what it is.
79
+ const again = deviceCause(second.out);
80
+ summary(`- \`${cmd.join(' ')}\` — **${again ? 'device unavailable' : 'check failed'}** after a revive${again ? ` (${again})` : ''}`);
81
+ if (again) console.error(`\nFAIL the simulator is still in a bad state after a revive: ${again}`);
82
+ process.exit(second.code ?? 1);
@@ -67,14 +67,17 @@ read_state() {
67
67
  printf '[{"button":"home"},{"settle":true}]\n' > /tmp/reset-local.json
68
68
  node src/cli.js do /tmp/reset-local.json --device="$DEVICE" >/dev/null 2>&1
69
69
  BEFORE=$(read_state)
70
- printf '[{"openUrl":"https://example.com"},{"settle":true}]\n' > /tmp/flow-local.json
71
- if node src/cli.js do /tmp/flow-local.json --device="$DEVICE" >/dev/null 2>&1; then
72
- AFTER=$(read_state)
73
- if [ "${BEFORE%% *}" != "${AFTER%% *}" ]; then ok "frame hash changed: ${BEFORE%% *} -> ${AFTER%% *}"
74
- else bad "a step ran cleanly but capture saw no change"; fi
75
- else
76
- bad "openUrl flow failed"
77
- fi
70
+ # From inside an app, pressing home always changes the screen — see the job's
71
+ # comment for the four vehicles that did not hold. Setup gets to a known screen;
72
+ # the asserted action is simframe's own HID path with no simctl in it.
73
+ printf '[{"launch":{"value":"com.apple.Preferences","relaunch":true}},{"settle":true}]\n' > /tmp/setup-local.json
74
+ node src/cli.js do /tmp/setup-local.json --device="$DEVICE" >/tmp/step-local.log 2>&1 || true
75
+ BEFORE=$(read_state)
76
+ printf '[{"button":"home"},{"settle":true}]\n' > /tmp/flow-local.json
77
+ node src/cli.js do /tmp/flow-local.json --device="$DEVICE" >>/tmp/step-local.log 2>&1 || true
78
+ AFTER=$(read_state)
79
+ if [ "${BEFORE%% *}" != "${AFTER%% *}" ]; then ok "frame hash changed: ${BEFORE%% *} -> ${AFTER%% *}"
80
+ else bad "a step ran and capture saw no change"; tail -5 /tmp/step-local.log; fi
78
81
 
79
82
  step "The memory layer — screen map, refs, graph, verdicts, flows"
80
83
  # Mirrors the job: exit 75 means the display wedged and nothing was tested, so
@@ -111,7 +111,14 @@ async function cli(args, { expectFail = false, allowFail = false } = {}) {
111
111
  // detail has been truncated for legibility, and the first version of this
112
112
  // guard looked for "did not produce a frame" in a string that had been cut
113
113
  // to "simframe daemon di". The full text only exists at this boundary.
114
- throw new Error(`simframe ${full.join(' ')} failed: ${why.slice(0, 400)}`);
114
+ //
115
+ // And when the payload is a JSON report, say what failed rather than
116
+ // handing back its first hundred characters. A run of this printed
117
+ // `simframe doctor --json failed: {\n "ok": false,\n "strict": true,\n
118
+ // "failu` — the word "failures" cut in half, one character before the only
119
+ // content that mattered. A harness that truncates away the reason is doing
120
+ // to its reader exactly what this repo keeps writing items about.
121
+ throw new Error(`simframe ${full.join(' ')} failed: ${summarise(why)}`);
115
122
  }
116
123
  }
117
124
 
@@ -193,6 +200,23 @@ async function jsonRetry(args, opts, attempts = 3) {
193
200
  throw last;
194
201
  }
195
202
 
203
+ /** A failed JSON report, reduced to the part that says what went wrong. */
204
+ function summarise(why) {
205
+ try {
206
+ const parsed = JSON.parse(why);
207
+ const failures = parsed.failures ?? parsed.failing ?? null;
208
+ if (Array.isArray(failures) && failures.length) {
209
+ return failures
210
+ .map((f) => (typeof f === 'string' ? f : `${f.name ?? f.check ?? '?'}: ${f.detail ?? f.note ?? f.message ?? ''}`.trim()))
211
+ .join('; ')
212
+ .slice(0, 400);
213
+ }
214
+ const bad = (parsed.checks ?? []).filter((c) => c.ok === false);
215
+ if (bad.length) return bad.map((c) => `${c.name}: ${c.detail ?? ''}`.trim()).join('; ').slice(0, 400);
216
+ } catch { /* not JSON, or not a shape we know — fall through to the raw text */ }
217
+ return why.slice(0, 400);
218
+ }
219
+
196
220
  const markHash = async () => (await jsonRetry(['mark'])).hash;
197
221
 
198
222
  function writeFlow(name, steps) {
@@ -201,10 +225,28 @@ function writeFlow(name, steps) {
201
225
  return file;
202
226
  }
203
227
 
204
- // A closed loop: openUrl puts Safari in front, home leaves it. Both ends are
205
- // screens the graph can learn, and every pass starts where the last one ended.
228
+ // A closed loop: launching an app puts it in front, home leaves it. Both ends
229
+ // are screens the graph can learn, and every pass starts where the last one
230
+ // ended.
231
+ //
232
+ // **It used to be `openUrl https://example.com`, and that was the bug.** Item
233
+ // 142 catalogued `simctl openurl` timing out on a loaded runner as a failure
234
+ // class, marked it "fixed — vehicle changed", and changed the vehicle in the
235
+ // *workflow's* step only. This loop was left on it, and so was the novel action
236
+ // below. That is the same class-versus-symptom error the `waitFor`/`assert`
237
+ // twin recorded: the fix went where the report pointed instead of everywhere
238
+ // the cause reached.
239
+ //
240
+ // It came back on 2026-09-14: passes 2 and 3 halted at step 0, the graph never
241
+ // got the chance to predict, and the failure read as "the outcome is predicted
242
+ // — pass 0", which names the graph for something Safari did.
243
+ //
244
+ // The replacement is the vehicle item 142 measured and proved for exactly this:
245
+ // **from inside an app, pressing home always changes the screen.** Settings is
246
+ // already installed everywhere this runs, the launch needs no network, and
247
+ // neither end depends on a browser cold-starting on a shared machine.
206
248
  const LOOP = writeFlow('simframe-ci-loop.json', [
207
- { openUrl: 'https://example.com' },
249
+ { launch: { value: 'com.apple.Preferences', relaunch: true } },
208
250
  { button: 'home' },
209
251
  ]);
210
252
  // Leaving whatever screen the map was read on.
@@ -513,10 +555,23 @@ if (novelRan && novelMoved) {
513
555
  // so a run in which every pass failed to dispatch says nothing about
514
556
  // prediction. It failed the build as `pass 0` while the real cause was a
515
557
  // simctl launch timing out, three checks upstream.
516
- const anyPassRan = passes.some((p) => Array.isArray(p.run?.results) && p.run.results.some((r) => r.ok !== false));
517
- if (!anyPassRan) {
558
+ //
559
+ // **The rule was right and the test of it was too coarse.** `anyPassRan` asks
560
+ // whether *any* pass ran, but prediction can only be observed on a pass AFTER
561
+ // the one that taught the edge — so pass 1 running is not enough. On
562
+ // 2026-09-14 pass 1 ran, passes 2 and 3 halted at step 0, and this reported
563
+ // `pass 0` as though the graph had declined to predict. Nothing had asked it
564
+ // to. Same sentence as the novel action three checks above: untested is not
565
+ // broken, and the guard has to test the pass the claim actually depends on.
566
+ const dispatched = (p) => Array.isArray(p.run?.results) && p.run.results.some((r) => r.ok !== false);
567
+ const laterPassRan = passes.slice(1).some(dispatched);
568
+ if (!passes.some(dispatched)) {
518
569
  skip('and once the graph has seen it, the outcome is predicted',
519
570
  'no pass dispatched a step, so the graph was never given anything to learn');
571
+ } else if (!laterPassRan) {
572
+ skip('and once the graph has seen it, the outcome is predicted',
573
+ `only the first pass dispatched a step (${passes.slice(1).map((p, i) => `pass ${i + 2}: [${p.verdicts.join(', ')}]`).join('; ')})`
574
+ + ' — prediction is only observable on a pass after the one that taught the edge');
520
575
  } else {
521
576
  check(passes.some((p) => p.verdicts.includes('ok')),
522
577
  'and once the graph has seen it, the outcome is predicted',
@@ -86,6 +86,19 @@ console.log(`tour: ${tour.length} screens x ${rounds} rounds, every reading cold
86
86
  const readings = [];
87
87
  /** Navigations that did not land before the reading was taken. */
88
88
  const arrivalFailures = [];
89
+ /** Readings too bare to be a screen — see the guard where this is used. */
90
+ const sparseReadings = [];
91
+ /** `name|round` of every reading that stayed too sparse, so the report below can name the cause. */
92
+ const sparseAt = new Set();
93
+ /**
94
+ * Below this, a reading cannot distinguish its screen from any other bare one.
95
+ *
96
+ * Not a failure threshold — see where it is used. The Settings root legitimately
97
+ * reads 4 tokens on a hosted runner, so treating this as "the screen has not
98
+ * drawn" rejected a real screen and made CI deterministically red. It marks
99
+ * readings worth suspecting when a score disagrees with itself, nothing more.
100
+ */
101
+ const MIN_TOKENS_FOR_A_READING = 5;
89
102
 
90
103
  /**
91
104
  * The tokens that carry a name, as opposed to a shape.
@@ -124,6 +137,76 @@ const save = (extra = {}) => {
124
137
  }, null, 2));
125
138
  };
126
139
 
140
+ /**
141
+ * How hard to try for a frame newer than the navigation before giving up.
142
+ *
143
+ * Three re-reads a second apart is ~3s of slack against capture medians that
144
+ * were measured above 1s on a bad runner. Generous enough to absorb the spikes
145
+ * that caused this, small enough that a genuinely stopped capture still fails
146
+ * rather than hanging the job.
147
+ */
148
+ const STALE_READ_RETRIES = 3;
149
+ /**
150
+ * How hard to try for a reading rich enough to tell its screen apart.
151
+ *
152
+ * Backed off rather than fixed, because the thing being waited for is a screen
153
+ * finishing its draw, and the runner that needs this is the slow one.
154
+ */
155
+ const SPARSE_READ_RETRIES = 3;
156
+ const SPARSE_READ_WAIT_MS = 1500;
157
+ const STALE_READ_WAIT_MS = 1000;
158
+
159
+ /** When the steps that were supposed to change the screen finished. */
160
+ let navigatedAt = 0;
161
+ /** Readings that never got a frame newer than their own navigation. */
162
+ const staleReadings = [];
163
+
164
+ /**
165
+ * Visit every screen once before measuring, taking no readings.
166
+ *
167
+ * **The rounds are supposed to be repeat measurements of one thing.** They were
168
+ * not: round 1 systematically differed from rounds 2 and 3 whenever an app was
169
+ * cold, because an app that has just launched has not finished publishing its
170
+ * accessibility tree. Measured on CI — round 1 of `browser` read `e00725fce7`
171
+ * with **0 named** elements where rounds 2 and 3 read `9589eb2471` with **5**,
172
+ * while every other screen matched across all three. That is the eval measuring
173
+ * app cold-start, which it never set out to measure and does not report.
174
+ *
175
+ * It was invisible for a long time because something else was paying for it:
176
+ * the memory-layer step runs first on CI and drives the same apps for 504s,
177
+ * Safari included, so the eval always met a warm device. Sharding the job
178
+ * removed that neighbour and the dependency surfaced immediately. The defect
179
+ * was always here; the neighbour was hiding it.
180
+ *
181
+ * This is **not** the same as making the readings warm. Every reading is still
182
+ * taken with `fresh: true` against cold screen memory, which is what "cold"
183
+ * means in this harness — the graph must not have seen the screen before. What
184
+ * the warm-up removes is a variable about the *operating system* that the
185
+ * fingerprint has nothing to do with.
186
+ *
187
+ * Skippable with `--no-warmup`, because the comparison is the evidence: run it
188
+ * both ways to see whether round 1 still disagrees with its own repeats.
189
+ */
190
+ async function warmUp() {
191
+ const started = Date.now();
192
+ console.log('warming: visiting each screen once, taking no readings');
193
+ for (const screen of tour) {
194
+ if (!screen.steps?.length) continue;
195
+ try {
196
+ await actions.runScript(device, { steps: screen.steps, verify: false });
197
+ } catch (err) {
198
+ // A warm-up failure is not a result. The measured rounds below will meet
199
+ // the same screen and fail there with the harness's own reporting, which
200
+ // says which screen and writes the readings out. Failing here would cost
201
+ // that and report a screen that was never measured.
202
+ console.log(` (warm-up could not reach "${screen.name}": ${err.message})`);
203
+ }
204
+ }
205
+ console.log(`warming: done in ${Math.round((Date.now() - started) / 1000)}s\n`);
206
+ }
207
+
208
+ if (!process.argv.includes('--no-warmup')) await warmUp();
209
+
127
210
  for (let round = 1; round <= rounds; round += 1) {
128
211
  for (const screen of tour) {
129
212
  if (screen.steps?.length) {
@@ -140,6 +223,7 @@ for (let round = 1; round <= rounds; round += 1) {
140
223
  // matters.
141
224
  try {
142
225
  await actions.runScript(device, { steps: screen.steps, verify: false });
226
+ navigatedAt = Date.now();
143
227
  } catch (err) {
144
228
  save({ abandonedAt: { screen: screen.name, round, error: err.message } });
145
229
  console.error(`\nFAIL round ${round}, "${screen.name}" never arrived: ${err.message}`);
@@ -148,7 +232,112 @@ for (let round = 1; round <= rounds; round += 1) {
148
232
  process.exit(1);
149
233
  }
150
234
  }
151
- const id = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
235
+ // A read can throw, and when it does the readings so far are the evidence.
236
+ //
237
+ // This file already says so — *"Everything read so far is written out
238
+ // first, because a failing run is the one whose evidence matters"* — and
239
+ // then left this call unguarded, so a runner whose OCR overran its budget
240
+ // produced a raw stack trace, no `--out` file, and nothing for
241
+ // `analyse-fingerprint.mjs` to read. Exactly the failure the `save()` above
242
+ // was written to prevent, one line away from it.
243
+ //
244
+ // Not retried here. The read budget is already the client's own give-up
245
+ // point, so a read that overran it is a statement about the machine, and
246
+ // the honest thing is to say which screen and how far the tour got.
247
+ let id;
248
+ try {
249
+ id = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
250
+ } catch (err) {
251
+ save({ abandonedAt: { screen: screen.name, round, error: err.message, phase: 'reading' } });
252
+ console.error(`\nFAIL round ${round}, "${screen.name}" could not be read: ${err.message}`);
253
+ console.error(`${readings.length} reading(s) taken before this were written out.`);
254
+ console.error('A read that overran its budget is a fact about the host, not about the');
255
+ console.error('fingerprint — the budget is already the client\'s own give-up point, so');
256
+ console.error('raising it would only move where the wait ends.');
257
+ process.exit(1);
258
+ }
259
+ // A reading off a frame older than the navigation is not a reading either.
260
+ //
261
+ // Same rule as the sparseness guard below, on the other axis, and it took a
262
+ // third symptom to see they were one cause. Three CI runs failed this step
263
+ // three different ways — `settings` reading 4 tokens, a reading that "does
264
+ // not resemble its own screen", and a reading "taken on the previous
265
+ // screen" at similarity 1.00 off a frame **7168ms old**. All three are the
266
+ // same sentence: the reading is not of the screen we think it is, because
267
+ // capture on a hosted runner is slow. The daemon log for that run shows
268
+ // capture medians of 417ms, 503ms, 996ms, 1106ms and 1132ms against the
269
+ // 55.4ms p50 measured for a healthy runner — 20x, and with damage-driven
270
+ // capture a `fresh` read returns the newest frame that EXISTS, which on a
271
+ // runner that far behind can predate the navigation entirely.
272
+ //
273
+ // So the frame must have been captured after the steps that were supposed
274
+ // to change the screen. Re-read rather than fail, and fail only if it stays
275
+ // stale — an eval that scores a stale frame is measuring the runner, which
276
+ // is the one thing this harness says it is not doing.
277
+ if (navigatedAt) {
278
+ for (let attempt = 0; attempt < STALE_READ_RETRIES; attempt += 1) {
279
+ const capturedAt = id.state?.capturedAt ?? 0;
280
+ if (capturedAt >= navigatedAt) break;
281
+ const age = Date.now() - capturedAt;
282
+ console.log(` ("${screen.name}" read a frame from ${age}ms ago, older than the navigation`
283
+ + ` — reading again ${attempt + 1}/${STALE_READ_RETRIES})`);
284
+ await new Promise((r) => setTimeout(r, STALE_READ_WAIT_MS));
285
+ id = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
286
+ }
287
+ if ((id.state?.capturedAt ?? 0) < navigatedAt) {
288
+ staleReadings.push(`${screen.name} round ${round}: frame still predates the navigation`
289
+ + ` by ${navigatedAt - (id.state?.capturedAt ?? 0)}ms after ${STALE_READ_RETRIES} re-reads`);
290
+ }
291
+ }
292
+ // A reading too sparse to be a screen is not a reading.
293
+ //
294
+ // On a runner measured at ~2.5x slower than a laptop (EXPERIMENTS §15),
295
+ // perception can land on a half-rendered screen: two CI runs recorded
296
+ // `settings` and `settings-general` **both at 4 tokens**, identical hash,
297
+ // and the eval duly reported that a reading did not resemble its own
298
+ // screen. It resembled nothing, because almost nothing had been drawn yet.
299
+ //
300
+ // This is the hazard `TOKEN_RULES_VERSION` 7 was written for, arriving
301
+ // through the harness instead of the rules: *"two sparse nameless readings
302
+ // then matched exactly, one hash standing for two different screens"*. A
303
+ // guard exists for identity and there was none here.
304
+ //
305
+ // Read again rather than fail, and fail only if it stays sparse — the same
306
+ // "untested is not passed" rule the memory harness learned. A screen that
307
+ // is genuinely this bare after a second look is a real finding.
308
+ if ((id.tokens ?? []).length < MIN_TOKENS_FOR_A_READING) {
309
+ // Read again, and keep the BEST reading rather than the last one.
310
+ //
311
+ // One retry was not enough and the failure it produced pointed at the
312
+ // wrong thing. On CI, `settings-general` read 4 tokens twice and then
313
+ // scored 1.00 against the *Settings root* — because the General screen's
314
+ // back button is labelled "Settings", so a General seen only down to its
315
+ // nav bar is structurally the same screen as the root. The harness
316
+ // reported it as a fingerprint that does not resemble its own screen. The
317
+ // fingerprint was fine; the look was too short. `settings-general` reads
318
+ // 5 tokens when it is read properly.
319
+ //
320
+ // Best-of, not last, because these reads are samples of a screen that is
321
+ // still drawing: a later read is usually richer but not reliably so, and
322
+ // throwing away a 5-token reading because the retry saw 4 would be the
323
+ // same bug with more steps.
324
+ const before = (id.tokens ?? []).length;
325
+ const seen = [before];
326
+ for (let attempt = 1; attempt <= SPARSE_READ_RETRIES; attempt += 1) {
327
+ await new Promise((r) => setTimeout(r, SPARSE_READ_WAIT_MS * attempt));
328
+ const again = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
329
+ seen.push((again.tokens ?? []).length);
330
+ if ((again.tokens ?? []).length > (id.tokens ?? []).length) id = again;
331
+ if ((id.tokens ?? []).length >= MIN_TOKENS_FOR_A_READING) break;
332
+ }
333
+ const after = (id.tokens ?? []).length;
334
+ console.log(` ("${screen.name}" read ${before} token(s) — too sparse to compare;`
335
+ + ` read again: ${seen.slice(1).join(', ')} — kept ${after})`);
336
+ if (after < MIN_TOKENS_FOR_A_READING) {
337
+ sparseReadings.push(`${screen.name} round ${round}: ${after} token(s) after ${seen.length} reads`);
338
+ sparseAt.add(`${screen.name}|${round}`);
339
+ }
340
+ }
152
341
  // Did we actually arrive? Two differently-named screens reading the same
153
342
  // fingerprint means the navigation did not land before the reading was
154
343
  // taken, and every distribution below it is then measuring the tour rather
@@ -219,9 +408,45 @@ for (let round = 1; round <= rounds; round += 1) {
219
408
  }
220
409
  }
221
410
 
222
- save();
411
+ save({ staleReadings });
223
412
  if (outFile) console.log(`\nwrote ${readings.length} readings to ${outFile}`);
224
413
 
414
+ if (sparseReadings.length) {
415
+ // Reported, NOT failed — and that correction is worth more than the check.
416
+ //
417
+ // This exited 1 for four hours on 2026-09-14 and turned an intermittent CI
418
+ // failure into a deterministic one. The reasoning was that a 4-token reading
419
+ // is a screen that has not drawn. It is not: the Settings root legitimately
420
+ // reads 4 tokens on a hosted runner, twice in a row, three times in a row —
421
+ // which the comment on `MIN_TOKENS_FOR_A_READING` had itself said ("4-8
422
+ // tokens") one screen above the code that rejected it.
423
+ //
424
+ // What remains true is that a reading this bare cannot distinguish its screen
425
+ // from anything else equally bare, so it is worth saying out loud. What is not
426
+ // true is that saying so should stop the run. The arrival check below is the
427
+ // one that catches the real collision, and it does so on evidence rather than
428
+ // on a token count.
429
+ console.log(`\nNOTE ${sparseReadings.length} reading(s) were too bare to distinguish a screen:`);
430
+ for (const f of sparseReadings) console.log(` ${f}`);
431
+ console.log(`\nFewer than ${MIN_TOKENS_FOR_A_READING} tokens after two reads. That is not automatically`);
432
+ console.log('wrong — a plain screen really can be this bare — but two readings this sparse');
433
+ console.log('cannot be told apart, which is the hazard TOKEN_RULES_VERSION 7 was written');
434
+ console.log('for. If a same-screen score below disagrees with itself, start here.');
435
+ }
436
+
437
+ if (staleReadings.length) {
438
+ // A NOTE and not a failure, for the same reason the sparseness guard is one:
439
+ // this says the runner was too slow to give a fresh frame, which is a fact
440
+ // about the machine and not about the fingerprint. It is printed *above* the
441
+ // arrival check on purpose — when both fire, this is the explanation of that.
442
+ console.log(`\nNOTE ${staleReadings.length} reading(s) never got a frame newer than their navigation:`);
443
+ for (const f of staleReadings) console.log(` ${f}`);
444
+ console.log(`\nCapture on this host is behind the tour. With damage-driven capture a "fresh"`);
445
+ console.log('read returns the newest frame that exists, so on a slow runner it can predate');
446
+ console.log('the navigation entirely. If an arrival failure follows, this is its cause and');
447
+ console.log('the tour is not what needs fixing.');
448
+ }
449
+
225
450
  if (arrivalFailures.length) {
226
451
  console.error(`\nFAIL ${arrivalFailures.length} reading(s) were taken on the previous screen:`);
227
452
  for (const f of arrivalFailures) console.error(` ${f}`);
@@ -296,6 +521,20 @@ if (strays.length) {
296
521
  const matchNamed = match ? namedTokens(match.tokens) : [];
297
522
  const collided = bestOther >= 0.99 && named.length === 0 && matchNamed.length === 0;
298
523
  if (collided) collisions += 1;
524
+ // The third cause, and the one that produced this report on 2026-09-15.
525
+ //
526
+ // A reading that stayed under the token floor did not fail to resemble its
527
+ // screen — it never saw enough of the screen to resemble anything. On CI
528
+ // `settings-general` read 4 tokens and scored 1.00 against the Settings
529
+ // root, because the General screen's back button is labelled "Settings", so
530
+ // a General seen only down to its nav bar IS the root structurally. The
531
+ // report called that a wrong turn and sent the reader to fix the tour.
532
+ //
533
+ // Not folded into `collided` above: that one means the fingerprint had
534
+ // nothing to work with, which is this harness's subject. This means we did
535
+ // not look long enough, which is the harness's own fault and a different
536
+ // remedy.
537
+ const underRead = sparseAt.has(`${reading.name}|${reading.round}`);
299
538
  console.error(` ${reading.name} r${reading.round}: own screen ${bestSelf.toFixed(2)}, `
300
539
  + `${match ? `${match.name} r${match.round}` : 'another screen'} ${bestOther.toFixed(2)} `
301
540
  + `(${reading.count} tokens, ${named.length} named, sources ${reading.sources.join('+') || 'none'})`);
@@ -303,6 +542,10 @@ if (strays.length) {
303
542
  console.error(' ^ a COLLISION, not a wrong turn: neither reading carries a chrome');
304
543
  console.error(' label, so both are structure with no name and the fingerprint has');
305
544
  console.error(' nothing left to tell two list screens apart.');
545
+ } else if (underRead) {
546
+ console.error(' ^ UNDER-READ, not a wrong turn: this reading stayed below the token');
547
+ console.error(' floor after every retry, so it never saw enough of its screen to');
548
+ console.error(' resemble one. Two screens read this thinly are the same screen.');
306
549
  }
307
550
  if (named.length) console.error(` names: ${named.map((t) => t.slice(t.indexOf('"'), t.lastIndexOf('"') + 1)).join(' ')}`);
308
551
  }
@@ -310,6 +553,10 @@ if (strays.length) {
310
553
  console.error(`\n${collisions} of ${strays.length} are fingerprint collisions. That is this harness's own subject,`);
311
554
  console.error('not a tour fault: a reading whose chrome label went missing cannot establish');
312
555
  console.error('identity, and comparing it as though it could is what produced the verdict above.');
556
+ } else if (strays.every((x) => sparseAt.has(`${x.reading.name}|${x.reading.round}`))) {
557
+ console.error('\nEvery stray above was UNDER-READ, so this says nothing about the tour or the');
558
+ console.error('fingerprint — the harness scored a look that was too short. The retries are in');
559
+ console.error('SPARSE_READ_RETRIES; a runner that needs more than they allow is the finding.');
313
560
  } else {
314
561
  console.error('\nThat is the tour going somewhere unintended, not the fingerprint drifting, and');
315
562
  console.error('measuring it as either distribution poisons both ends. Fix the tour — a tap that');