simframe 0.13.0 → 0.14.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +15 -0
- package/native/simframed/Sources/SimframeCore/Motion.swift +33 -2
- package/package.json +1 -1
- package/scripts/ci-device-guard.mjs +82 -0
- package/scripts/ci-integration-local.sh +11 -8
- package/scripts/ci-memory.mjs +61 -6
- package/scripts/eval-fingerprint.mjs +249 -2
- package/src/actions.js +117 -4
- package/src/cli.js +23 -2
- package/src/graph.js +15 -2
- package/src/matching.js +17 -1
- package/src/mcp.js +141 -12
- package/src/navigate.js +4 -1
- package/src/platform/android.js +34 -0
- package/src/platform/index.js +7 -0
- package/src/platform/ios.js +158 -8
- package/src/platform/plist.js +156 -0
- package/src/storage.js +201 -0
package/README.md
CHANGED
|
@@ -406,6 +406,7 @@ steer the model is a tool surface the model uses wrong.
|
|
|
406
406
|
| `sim_wait` | Waits for the screen to change *and then* settle. |
|
|
407
407
|
| `sim_look` | **The only tool that returns an image**, capped at 1024 px. For layout, colour, spacing — questions text cannot answer. |
|
|
408
408
|
| `sim_recall` · `sim_strip` | Look backwards: a text timeline of what happened, or recent frames tiled into one image. |
|
|
409
|
+
| `sim_storage` | **What the app believes**, as opposed to what it drew: its `UserDefaults` and, for React Native, its `AsyncStorage`. Reads the data container off disk, so it answers on a device that is **not running**. |
|
|
409
410
|
| `sim_capture` · `sim_devices` | Manage capture loops; list simulators. |
|
|
410
411
|
|
|
411
412
|
### What the screen looks like as text
|
|
@@ -585,6 +586,20 @@ simframe screens # what this device has learned
|
|
|
585
586
|
simframe goto invoices # walk there, verifying every step
|
|
586
587
|
```
|
|
587
588
|
|
|
589
|
+
And when the screen and the behaviour disagree, the question is usually not
|
|
590
|
+
about the screen at all:
|
|
591
|
+
|
|
592
|
+
```bash
|
|
593
|
+
simframe storage # apps with a data container
|
|
594
|
+
simframe storage com.example.myapp # what that app saved
|
|
595
|
+
```
|
|
596
|
+
|
|
597
|
+
`sim_ui` says what is drawn; `sim_storage` says what the app believes. It reads
|
|
598
|
+
the data container straight off the host filesystem, which means it works on a
|
|
599
|
+
device that is **shut down** — `simctl` cannot do this at all, on any of its own
|
|
600
|
+
paths, once a device stops running.
|
|
601
|
+
|
|
602
|
+
|
|
588
603
|
Measured on a four-tab tour, `goto` plans and walks three-step routes with every
|
|
589
604
|
step verified and no model call. It fails rather than guesses: an unknown
|
|
590
605
|
destination, a query matching two screens equally, or no path of known edges all
|
|
@@ -13,6 +13,10 @@ public enum Motion {
|
|
|
13
13
|
|
|
14
14
|
/// Below this mean absolute difference, two frames are the same picture.
|
|
15
15
|
public static let stillThreshold = 0.004
|
|
16
|
+
|
|
17
|
+
/// How many moving cells make an animation rather than sensor noise.
|
|
18
|
+
/// Set from the measurements in `state(history:now:)` below.
|
|
19
|
+
public static let minAnimatingCells = 12
|
|
16
20
|
/// Frames that must agree before the screen counts as settled.
|
|
17
21
|
///
|
|
18
22
|
/// Paired with a duration, because frame count alone is not a measure of
|
|
@@ -170,8 +174,35 @@ public enum Motion {
|
|
|
170
174
|
let cellThreshold = 0.06
|
|
171
175
|
let moving = union.filter { $0 > cellThreshold }.count
|
|
172
176
|
let fraction = Double(moving) / Double(union.count)
|
|
173
|
-
// Small and persistent, rather than a screen changing
|
|
174
|
-
|
|
177
|
+
// Small and persistent, rather than a screen changing — and big
|
|
178
|
+
// enough to be something.
|
|
179
|
+
//
|
|
180
|
+
// `fraction > 0` meant one cell of 4,608 counted as an animation,
|
|
181
|
+
// and measured on a device that fires constantly on screens where
|
|
182
|
+
// nothing is happening. A field report put it exactly right: a
|
|
183
|
+
// warning that is usually wrong trains the reader to ignore the one
|
|
184
|
+
// that matters.
|
|
185
|
+
//
|
|
186
|
+
// Measured on this device, 2026-09-14, which is what
|
|
187
|
+
// `minAnimatingCells` is set from:
|
|
188
|
+
//
|
|
189
|
+
// static home screen, nothing moving 1-4 cells, on 54% of frames
|
|
190
|
+
// a blinking text caret **2 cells** (1x2), on 100%
|
|
191
|
+
// the testbed's spinner 169 cells (13x13)
|
|
192
|
+
// Maps launching, median 656 cells (41x16)
|
|
193
|
+
// a real spinner, from the field 1,364 cells (44x31)
|
|
194
|
+
//
|
|
195
|
+
// Caret and noise sit at 1-4; the smallest real animation seen is
|
|
196
|
+
// 169. A 42x gap with nothing in it, so the threshold is not
|
|
197
|
+
// delicate — 12 is three times the worst noise and an order of
|
|
198
|
+
// magnitude below the weakest signal.
|
|
199
|
+
//
|
|
200
|
+
// The caret is the case this is really for, and it is why the
|
|
201
|
+
// question could not be settled by reasoning: a caret must never
|
|
202
|
+
// stop a screen from settling, and before this it did — 72 frames
|
|
203
|
+
// out of 72 with `settled: false` on a screen holding nothing but a
|
|
204
|
+
// text cursor.
|
|
205
|
+
if moving >= Self.minAnimatingCells && fraction < 0.06 {
|
|
175
206
|
animating = boundingBox(of: union, threshold: cellThreshold)
|
|
176
207
|
}
|
|
177
208
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.14.1",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
5
|
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Run a CI step, and tell a sick simulator apart from a failing check.
|
|
3
|
+
//
|
|
4
|
+
// node scripts/ci-device-guard.mjs <udid> -- <command> [args...]
|
|
5
|
+
//
|
|
6
|
+
// **Why this exists, with the number that justifies it.** Over the last 25 CI
|
|
7
|
+
// runs the `integration` job failed 17 times and passed 5. Eight of the nine
|
|
8
|
+
// most recent failures name a device-state condition in their own output —
|
|
9
|
+
// `NSPOSIXErrorDomain code=60`, `the display produced no frame in 60s`, `the
|
|
10
|
+
// second app never launched` — and one of those runs was a **docs-only commit**
|
|
11
|
+
// that changed a single markdown file. Two runs of byte-identical code failed at
|
|
12
|
+
// two different steps.
|
|
13
|
+
//
|
|
14
|
+
// So the job has been answering two questions at once — *does simframe work* and
|
|
15
|
+
// *did this hosted simulator survive twenty minutes* — and the second dominates.
|
|
16
|
+
// A red build that is usually the device is the cry-wolf failure this repo keeps
|
|
17
|
+
// writing items about, and it cost two days of re-reading logs to learn nothing.
|
|
18
|
+
//
|
|
19
|
+
// This does not paper over failures. It distinguishes them: a named device
|
|
20
|
+
// condition gets the cure this project already ships (`simframe revive`) and one
|
|
21
|
+
// retry, exactly as DEFERRED 126 says to; anything else fails on the spot,
|
|
22
|
+
// untouched. The classification is written to the job summary so the *rate*
|
|
23
|
+
// becomes visible instead of arguable.
|
|
24
|
+
import { spawn } from 'node:child_process';
|
|
25
|
+
import fs from 'node:fs';
|
|
26
|
+
|
|
27
|
+
/** Conditions that are the simulator, not the code. Each seen in a real run. */
|
|
28
|
+
const DEVICE_STATE = [
|
|
29
|
+
[/NSPOSIXErrorDomain.*code=?\s*60|Operation timed out/i, 'simctl stopped answering (NSPOSIXErrorDomain 60)'],
|
|
30
|
+
[/did not produce a frame|produced no frame in \d+s/i, 'the daemon is up and the display renders nothing'],
|
|
31
|
+
[/Timeout waiting for screen surfaces|display surface is not answering|display surface could not be read/i, 'the display surface is wedged'],
|
|
32
|
+
[/no frames buffered|capture is wedged/i, 'capture stopped'],
|
|
33
|
+
[/the second app never launched|could not be dispatched/i, 'an app would not launch'],
|
|
34
|
+
];
|
|
35
|
+
|
|
36
|
+
const udid = process.argv[2];
|
|
37
|
+
const sep = process.argv.indexOf('--');
|
|
38
|
+
if (!udid || sep < 0) {
|
|
39
|
+
console.error('usage: ci-device-guard.mjs <udid> -- <command> [args...]');
|
|
40
|
+
process.exit(2);
|
|
41
|
+
}
|
|
42
|
+
const cmd = process.argv.slice(sep + 1);
|
|
43
|
+
|
|
44
|
+
function run(argv, { capture = true } = {}) {
|
|
45
|
+
return new Promise((resolve) => {
|
|
46
|
+
const p = spawn(argv[0], argv.slice(1), { stdio: capture ? ['inherit', 'pipe', 'pipe'] : 'inherit' });
|
|
47
|
+
let out = '';
|
|
48
|
+
p.stdout?.on('data', (d) => { out += d; process.stdout.write(d); });
|
|
49
|
+
p.stderr?.on('data', (d) => { out += d; process.stderr.write(d); });
|
|
50
|
+
p.on('close', (code) => resolve({ code, out }));
|
|
51
|
+
});
|
|
52
|
+
}
|
|
53
|
+
|
|
54
|
+
const summary = (line) => {
|
|
55
|
+
const f = process.env.GITHUB_STEP_SUMMARY;
|
|
56
|
+
if (f) { try { fs.appendFileSync(f, `${line}\n`); } catch { /* summaries are a nicety */ } }
|
|
57
|
+
};
|
|
58
|
+
|
|
59
|
+
const deviceCause = (text) => DEVICE_STATE.find(([re]) => re.test(text))?.[1] ?? null;
|
|
60
|
+
|
|
61
|
+
const first = await run(cmd);
|
|
62
|
+
if (first.code === 0) process.exit(0);
|
|
63
|
+
|
|
64
|
+
const cause = deviceCause(first.out);
|
|
65
|
+
if (!cause) {
|
|
66
|
+
console.error(`\n (this step failed on its merits, not on the device — not retrying)`);
|
|
67
|
+
summary(`- \`${cmd.join(' ')}\` — **check failed** (exit ${first.code})`);
|
|
68
|
+
process.exit(first.code ?? 1);
|
|
69
|
+
}
|
|
70
|
+
|
|
71
|
+
console.error(`\n (${cause} — DEFERRED 126. Reviving once and running again.)`);
|
|
72
|
+
await run(['node', 'src/cli.js', 'revive', `--device=${udid}`], { capture: false });
|
|
73
|
+
const second = await run(cmd);
|
|
74
|
+
if (second.code === 0) {
|
|
75
|
+
summary(`- \`${cmd.join(' ')}\` — passed after one revive (${cause})`);
|
|
76
|
+
process.exit(0);
|
|
77
|
+
}
|
|
78
|
+
// Twice in a row, on a condition we know the cure for. Reported as what it is.
|
|
79
|
+
const again = deviceCause(second.out);
|
|
80
|
+
summary(`- \`${cmd.join(' ')}\` — **${again ? 'device unavailable' : 'check failed'}** after a revive${again ? ` (${again})` : ''}`);
|
|
81
|
+
if (again) console.error(`\nFAIL the simulator is still in a bad state after a revive: ${again}`);
|
|
82
|
+
process.exit(second.code ?? 1);
|
|
@@ -67,14 +67,17 @@ read_state() {
|
|
|
67
67
|
printf '[{"button":"home"},{"settle":true}]\n' > /tmp/reset-local.json
|
|
68
68
|
node src/cli.js do /tmp/reset-local.json --device="$DEVICE" >/dev/null 2>&1
|
|
69
69
|
BEFORE=$(read_state)
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
70
|
+
# From inside an app, pressing home always changes the screen — see the job's
|
|
71
|
+
# comment for the four vehicles that did not hold. Setup gets to a known screen;
|
|
72
|
+
# the asserted action is simframe's own HID path with no simctl in it.
|
|
73
|
+
printf '[{"launch":{"value":"com.apple.Preferences","relaunch":true}},{"settle":true}]\n' > /tmp/setup-local.json
|
|
74
|
+
node src/cli.js do /tmp/setup-local.json --device="$DEVICE" >/tmp/step-local.log 2>&1 || true
|
|
75
|
+
BEFORE=$(read_state)
|
|
76
|
+
printf '[{"button":"home"},{"settle":true}]\n' > /tmp/flow-local.json
|
|
77
|
+
node src/cli.js do /tmp/flow-local.json --device="$DEVICE" >>/tmp/step-local.log 2>&1 || true
|
|
78
|
+
AFTER=$(read_state)
|
|
79
|
+
if [ "${BEFORE%% *}" != "${AFTER%% *}" ]; then ok "frame hash changed: ${BEFORE%% *} -> ${AFTER%% *}"
|
|
80
|
+
else bad "a step ran and capture saw no change"; tail -5 /tmp/step-local.log; fi
|
|
78
81
|
|
|
79
82
|
step "The memory layer — screen map, refs, graph, verdicts, flows"
|
|
80
83
|
# Mirrors the job: exit 75 means the display wedged and nothing was tested, so
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -111,7 +111,14 @@ async function cli(args, { expectFail = false, allowFail = false } = {}) {
|
|
|
111
111
|
// detail has been truncated for legibility, and the first version of this
|
|
112
112
|
// guard looked for "did not produce a frame" in a string that had been cut
|
|
113
113
|
// to "simframe daemon di". The full text only exists at this boundary.
|
|
114
|
-
|
|
114
|
+
//
|
|
115
|
+
// And when the payload is a JSON report, say what failed rather than
|
|
116
|
+
// handing back its first hundred characters. A run of this printed
|
|
117
|
+
// `simframe doctor --json failed: {\n "ok": false,\n "strict": true,\n
|
|
118
|
+
// "failu` — the word "failures" cut in half, one character before the only
|
|
119
|
+
// content that mattered. A harness that truncates away the reason is doing
|
|
120
|
+
// to its reader exactly what this repo keeps writing items about.
|
|
121
|
+
throw new Error(`simframe ${full.join(' ')} failed: ${summarise(why)}`);
|
|
115
122
|
}
|
|
116
123
|
}
|
|
117
124
|
|
|
@@ -193,6 +200,23 @@ async function jsonRetry(args, opts, attempts = 3) {
|
|
|
193
200
|
throw last;
|
|
194
201
|
}
|
|
195
202
|
|
|
203
|
+
/** A failed JSON report, reduced to the part that says what went wrong. */
|
|
204
|
+
function summarise(why) {
|
|
205
|
+
try {
|
|
206
|
+
const parsed = JSON.parse(why);
|
|
207
|
+
const failures = parsed.failures ?? parsed.failing ?? null;
|
|
208
|
+
if (Array.isArray(failures) && failures.length) {
|
|
209
|
+
return failures
|
|
210
|
+
.map((f) => (typeof f === 'string' ? f : `${f.name ?? f.check ?? '?'}: ${f.detail ?? f.note ?? f.message ?? ''}`.trim()))
|
|
211
|
+
.join('; ')
|
|
212
|
+
.slice(0, 400);
|
|
213
|
+
}
|
|
214
|
+
const bad = (parsed.checks ?? []).filter((c) => c.ok === false);
|
|
215
|
+
if (bad.length) return bad.map((c) => `${c.name}: ${c.detail ?? ''}`.trim()).join('; ').slice(0, 400);
|
|
216
|
+
} catch { /* not JSON, or not a shape we know — fall through to the raw text */ }
|
|
217
|
+
return why.slice(0, 400);
|
|
218
|
+
}
|
|
219
|
+
|
|
196
220
|
const markHash = async () => (await jsonRetry(['mark'])).hash;
|
|
197
221
|
|
|
198
222
|
function writeFlow(name, steps) {
|
|
@@ -201,10 +225,28 @@ function writeFlow(name, steps) {
|
|
|
201
225
|
return file;
|
|
202
226
|
}
|
|
203
227
|
|
|
204
|
-
// A closed loop:
|
|
205
|
-
// screens the graph can learn, and every pass starts where the last one
|
|
228
|
+
// A closed loop: launching an app puts it in front, home leaves it. Both ends
|
|
229
|
+
// are screens the graph can learn, and every pass starts where the last one
|
|
230
|
+
// ended.
|
|
231
|
+
//
|
|
232
|
+
// **It used to be `openUrl https://example.com`, and that was the bug.** Item
|
|
233
|
+
// 142 catalogued `simctl openurl` timing out on a loaded runner as a failure
|
|
234
|
+
// class, marked it "fixed — vehicle changed", and changed the vehicle in the
|
|
235
|
+
// *workflow's* step only. This loop was left on it, and so was the novel action
|
|
236
|
+
// below. That is the same class-versus-symptom error the `waitFor`/`assert`
|
|
237
|
+
// twin recorded: the fix went where the report pointed instead of everywhere
|
|
238
|
+
// the cause reached.
|
|
239
|
+
//
|
|
240
|
+
// It came back on 2026-09-14: passes 2 and 3 halted at step 0, the graph never
|
|
241
|
+
// got the chance to predict, and the failure read as "the outcome is predicted
|
|
242
|
+
// — pass 0", which names the graph for something Safari did.
|
|
243
|
+
//
|
|
244
|
+
// The replacement is the vehicle item 142 measured and proved for exactly this:
|
|
245
|
+
// **from inside an app, pressing home always changes the screen.** Settings is
|
|
246
|
+
// already installed everywhere this runs, the launch needs no network, and
|
|
247
|
+
// neither end depends on a browser cold-starting on a shared machine.
|
|
206
248
|
const LOOP = writeFlow('simframe-ci-loop.json', [
|
|
207
|
-
{
|
|
249
|
+
{ launch: { value: 'com.apple.Preferences', relaunch: true } },
|
|
208
250
|
{ button: 'home' },
|
|
209
251
|
]);
|
|
210
252
|
// Leaving whatever screen the map was read on.
|
|
@@ -513,10 +555,23 @@ if (novelRan && novelMoved) {
|
|
|
513
555
|
// so a run in which every pass failed to dispatch says nothing about
|
|
514
556
|
// prediction. It failed the build as `pass 0` while the real cause was a
|
|
515
557
|
// simctl launch timing out, three checks upstream.
|
|
516
|
-
|
|
517
|
-
|
|
558
|
+
//
|
|
559
|
+
// **The rule was right and the test of it was too coarse.** `anyPassRan` asks
|
|
560
|
+
// whether *any* pass ran, but prediction can only be observed on a pass AFTER
|
|
561
|
+
// the one that taught the edge — so pass 1 running is not enough. On
|
|
562
|
+
// 2026-09-14 pass 1 ran, passes 2 and 3 halted at step 0, and this reported
|
|
563
|
+
// `pass 0` as though the graph had declined to predict. Nothing had asked it
|
|
564
|
+
// to. Same sentence as the novel action three checks above: untested is not
|
|
565
|
+
// broken, and the guard has to test the pass the claim actually depends on.
|
|
566
|
+
const dispatched = (p) => Array.isArray(p.run?.results) && p.run.results.some((r) => r.ok !== false);
|
|
567
|
+
const laterPassRan = passes.slice(1).some(dispatched);
|
|
568
|
+
if (!passes.some(dispatched)) {
|
|
518
569
|
skip('and once the graph has seen it, the outcome is predicted',
|
|
519
570
|
'no pass dispatched a step, so the graph was never given anything to learn');
|
|
571
|
+
} else if (!laterPassRan) {
|
|
572
|
+
skip('and once the graph has seen it, the outcome is predicted',
|
|
573
|
+
`only the first pass dispatched a step (${passes.slice(1).map((p, i) => `pass ${i + 2}: [${p.verdicts.join(', ')}]`).join('; ')})`
|
|
574
|
+
+ ' — prediction is only observable on a pass after the one that taught the edge');
|
|
520
575
|
} else {
|
|
521
576
|
check(passes.some((p) => p.verdicts.includes('ok')),
|
|
522
577
|
'and once the graph has seen it, the outcome is predicted',
|
|
@@ -86,6 +86,19 @@ console.log(`tour: ${tour.length} screens x ${rounds} rounds, every reading cold
|
|
|
86
86
|
const readings = [];
|
|
87
87
|
/** Navigations that did not land before the reading was taken. */
|
|
88
88
|
const arrivalFailures = [];
|
|
89
|
+
/** Readings too bare to be a screen — see the guard where this is used. */
|
|
90
|
+
const sparseReadings = [];
|
|
91
|
+
/** `name|round` of every reading that stayed too sparse, so the report below can name the cause. */
|
|
92
|
+
const sparseAt = new Set();
|
|
93
|
+
/**
|
|
94
|
+
* Below this, a reading cannot distinguish its screen from any other bare one.
|
|
95
|
+
*
|
|
96
|
+
* Not a failure threshold — see where it is used. The Settings root legitimately
|
|
97
|
+
* reads 4 tokens on a hosted runner, so treating this as "the screen has not
|
|
98
|
+
* drawn" rejected a real screen and made CI deterministically red. It marks
|
|
99
|
+
* readings worth suspecting when a score disagrees with itself, nothing more.
|
|
100
|
+
*/
|
|
101
|
+
const MIN_TOKENS_FOR_A_READING = 5;
|
|
89
102
|
|
|
90
103
|
/**
|
|
91
104
|
* The tokens that carry a name, as opposed to a shape.
|
|
@@ -124,6 +137,76 @@ const save = (extra = {}) => {
|
|
|
124
137
|
}, null, 2));
|
|
125
138
|
};
|
|
126
139
|
|
|
140
|
+
/**
|
|
141
|
+
* How hard to try for a frame newer than the navigation before giving up.
|
|
142
|
+
*
|
|
143
|
+
* Three re-reads a second apart is ~3s of slack against capture medians that
|
|
144
|
+
* were measured above 1s on a bad runner. Generous enough to absorb the spikes
|
|
145
|
+
* that caused this, small enough that a genuinely stopped capture still fails
|
|
146
|
+
* rather than hanging the job.
|
|
147
|
+
*/
|
|
148
|
+
const STALE_READ_RETRIES = 3;
|
|
149
|
+
/**
|
|
150
|
+
* How hard to try for a reading rich enough to tell its screen apart.
|
|
151
|
+
*
|
|
152
|
+
* Backed off rather than fixed, because the thing being waited for is a screen
|
|
153
|
+
* finishing its draw, and the runner that needs this is the slow one.
|
|
154
|
+
*/
|
|
155
|
+
const SPARSE_READ_RETRIES = 3;
|
|
156
|
+
const SPARSE_READ_WAIT_MS = 1500;
|
|
157
|
+
const STALE_READ_WAIT_MS = 1000;
|
|
158
|
+
|
|
159
|
+
/** When the steps that were supposed to change the screen finished. */
|
|
160
|
+
let navigatedAt = 0;
|
|
161
|
+
/** Readings that never got a frame newer than their own navigation. */
|
|
162
|
+
const staleReadings = [];
|
|
163
|
+
|
|
164
|
+
/**
|
|
165
|
+
* Visit every screen once before measuring, taking no readings.
|
|
166
|
+
*
|
|
167
|
+
* **The rounds are supposed to be repeat measurements of one thing.** They were
|
|
168
|
+
* not: round 1 systematically differed from rounds 2 and 3 whenever an app was
|
|
169
|
+
* cold, because an app that has just launched has not finished publishing its
|
|
170
|
+
* accessibility tree. Measured on CI — round 1 of `browser` read `e00725fce7`
|
|
171
|
+
* with **0 named** elements where rounds 2 and 3 read `9589eb2471` with **5**,
|
|
172
|
+
* while every other screen matched across all three. That is the eval measuring
|
|
173
|
+
* app cold-start, which it never set out to measure and does not report.
|
|
174
|
+
*
|
|
175
|
+
* It was invisible for a long time because something else was paying for it:
|
|
176
|
+
* the memory-layer step runs first on CI and drives the same apps for 504s,
|
|
177
|
+
* Safari included, so the eval always met a warm device. Sharding the job
|
|
178
|
+
* removed that neighbour and the dependency surfaced immediately. The defect
|
|
179
|
+
* was always here; the neighbour was hiding it.
|
|
180
|
+
*
|
|
181
|
+
* This is **not** the same as making the readings warm. Every reading is still
|
|
182
|
+
* taken with `fresh: true` against cold screen memory, which is what "cold"
|
|
183
|
+
* means in this harness — the graph must not have seen the screen before. What
|
|
184
|
+
* the warm-up removes is a variable about the *operating system* that the
|
|
185
|
+
* fingerprint has nothing to do with.
|
|
186
|
+
*
|
|
187
|
+
* Skippable with `--no-warmup`, because the comparison is the evidence: run it
|
|
188
|
+
* both ways to see whether round 1 still disagrees with its own repeats.
|
|
189
|
+
*/
|
|
190
|
+
async function warmUp() {
|
|
191
|
+
const started = Date.now();
|
|
192
|
+
console.log('warming: visiting each screen once, taking no readings');
|
|
193
|
+
for (const screen of tour) {
|
|
194
|
+
if (!screen.steps?.length) continue;
|
|
195
|
+
try {
|
|
196
|
+
await actions.runScript(device, { steps: screen.steps, verify: false });
|
|
197
|
+
} catch (err) {
|
|
198
|
+
// A warm-up failure is not a result. The measured rounds below will meet
|
|
199
|
+
// the same screen and fail there with the harness's own reporting, which
|
|
200
|
+
// says which screen and writes the readings out. Failing here would cost
|
|
201
|
+
// that and report a screen that was never measured.
|
|
202
|
+
console.log(` (warm-up could not reach "${screen.name}": ${err.message})`);
|
|
203
|
+
}
|
|
204
|
+
}
|
|
205
|
+
console.log(`warming: done in ${Math.round((Date.now() - started) / 1000)}s\n`);
|
|
206
|
+
}
|
|
207
|
+
|
|
208
|
+
if (!process.argv.includes('--no-warmup')) await warmUp();
|
|
209
|
+
|
|
127
210
|
for (let round = 1; round <= rounds; round += 1) {
|
|
128
211
|
for (const screen of tour) {
|
|
129
212
|
if (screen.steps?.length) {
|
|
@@ -140,6 +223,7 @@ for (let round = 1; round <= rounds; round += 1) {
|
|
|
140
223
|
// matters.
|
|
141
224
|
try {
|
|
142
225
|
await actions.runScript(device, { steps: screen.steps, verify: false });
|
|
226
|
+
navigatedAt = Date.now();
|
|
143
227
|
} catch (err) {
|
|
144
228
|
save({ abandonedAt: { screen: screen.name, round, error: err.message } });
|
|
145
229
|
console.error(`\nFAIL round ${round}, "${screen.name}" never arrived: ${err.message}`);
|
|
@@ -148,7 +232,112 @@ for (let round = 1; round <= rounds; round += 1) {
|
|
|
148
232
|
process.exit(1);
|
|
149
233
|
}
|
|
150
234
|
}
|
|
151
|
-
|
|
235
|
+
// A read can throw, and when it does the readings so far are the evidence.
|
|
236
|
+
//
|
|
237
|
+
// This file already says so — *"Everything read so far is written out
|
|
238
|
+
// first, because a failing run is the one whose evidence matters"* — and
|
|
239
|
+
// then left this call unguarded, so a runner whose OCR overran its budget
|
|
240
|
+
// produced a raw stack trace, no `--out` file, and nothing for
|
|
241
|
+
// `analyse-fingerprint.mjs` to read. Exactly the failure the `save()` above
|
|
242
|
+
// was written to prevent, one line away from it.
|
|
243
|
+
//
|
|
244
|
+
// Not retried here. The read budget is already the client's own give-up
|
|
245
|
+
// point, so a read that overran it is a statement about the machine, and
|
|
246
|
+
// the honest thing is to say which screen and how far the tour got.
|
|
247
|
+
let id;
|
|
248
|
+
try {
|
|
249
|
+
id = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
|
|
250
|
+
} catch (err) {
|
|
251
|
+
save({ abandonedAt: { screen: screen.name, round, error: err.message, phase: 'reading' } });
|
|
252
|
+
console.error(`\nFAIL round ${round}, "${screen.name}" could not be read: ${err.message}`);
|
|
253
|
+
console.error(`${readings.length} reading(s) taken before this were written out.`);
|
|
254
|
+
console.error('A read that overran its budget is a fact about the host, not about the');
|
|
255
|
+
console.error('fingerprint — the budget is already the client\'s own give-up point, so');
|
|
256
|
+
console.error('raising it would only move where the wait ends.');
|
|
257
|
+
process.exit(1);
|
|
258
|
+
}
|
|
259
|
+
// A reading off a frame older than the navigation is not a reading either.
|
|
260
|
+
//
|
|
261
|
+
// Same rule as the sparseness guard below, on the other axis, and it took a
|
|
262
|
+
// third symptom to see they were one cause. Three CI runs failed this step
|
|
263
|
+
// three different ways — `settings` reading 4 tokens, a reading that "does
|
|
264
|
+
// not resemble its own screen", and a reading "taken on the previous
|
|
265
|
+
// screen" at similarity 1.00 off a frame **7168ms old**. All three are the
|
|
266
|
+
// same sentence: the reading is not of the screen we think it is, because
|
|
267
|
+
// capture on a hosted runner is slow. The daemon log for that run shows
|
|
268
|
+
// capture medians of 417ms, 503ms, 996ms, 1106ms and 1132ms against the
|
|
269
|
+
// 55.4ms p50 measured for a healthy runner — 20x, and with damage-driven
|
|
270
|
+
// capture a `fresh` read returns the newest frame that EXISTS, which on a
|
|
271
|
+
// runner that far behind can predate the navigation entirely.
|
|
272
|
+
//
|
|
273
|
+
// So the frame must have been captured after the steps that were supposed
|
|
274
|
+
// to change the screen. Re-read rather than fail, and fail only if it stays
|
|
275
|
+
// stale — an eval that scores a stale frame is measuring the runner, which
|
|
276
|
+
// is the one thing this harness says it is not doing.
|
|
277
|
+
if (navigatedAt) {
|
|
278
|
+
for (let attempt = 0; attempt < STALE_READ_RETRIES; attempt += 1) {
|
|
279
|
+
const capturedAt = id.state?.capturedAt ?? 0;
|
|
280
|
+
if (capturedAt >= navigatedAt) break;
|
|
281
|
+
const age = Date.now() - capturedAt;
|
|
282
|
+
console.log(` ("${screen.name}" read a frame from ${age}ms ago, older than the navigation`
|
|
283
|
+
+ ` — reading again ${attempt + 1}/${STALE_READ_RETRIES})`);
|
|
284
|
+
await new Promise((r) => setTimeout(r, STALE_READ_WAIT_MS));
|
|
285
|
+
id = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
|
|
286
|
+
}
|
|
287
|
+
if ((id.state?.capturedAt ?? 0) < navigatedAt) {
|
|
288
|
+
staleReadings.push(`${screen.name} round ${round}: frame still predates the navigation`
|
|
289
|
+
+ ` by ${navigatedAt - (id.state?.capturedAt ?? 0)}ms after ${STALE_READ_RETRIES} re-reads`);
|
|
290
|
+
}
|
|
291
|
+
}
|
|
292
|
+
// A reading too sparse to be a screen is not a reading.
|
|
293
|
+
//
|
|
294
|
+
// On a runner measured at ~2.5x slower than a laptop (EXPERIMENTS §15),
|
|
295
|
+
// perception can land on a half-rendered screen: two CI runs recorded
|
|
296
|
+
// `settings` and `settings-general` **both at 4 tokens**, identical hash,
|
|
297
|
+
// and the eval duly reported that a reading did not resemble its own
|
|
298
|
+
// screen. It resembled nothing, because almost nothing had been drawn yet.
|
|
299
|
+
//
|
|
300
|
+
// This is the hazard `TOKEN_RULES_VERSION` 7 was written for, arriving
|
|
301
|
+
// through the harness instead of the rules: *"two sparse nameless readings
|
|
302
|
+
// then matched exactly, one hash standing for two different screens"*. A
|
|
303
|
+
// guard exists for identity and there was none here.
|
|
304
|
+
//
|
|
305
|
+
// Read again rather than fail, and fail only if it stays sparse — the same
|
|
306
|
+
// "untested is not passed" rule the memory harness learned. A screen that
|
|
307
|
+
// is genuinely this bare after a second look is a real finding.
|
|
308
|
+
if ((id.tokens ?? []).length < MIN_TOKENS_FOR_A_READING) {
|
|
309
|
+
// Read again, and keep the BEST reading rather than the last one.
|
|
310
|
+
//
|
|
311
|
+
// One retry was not enough and the failure it produced pointed at the
|
|
312
|
+
// wrong thing. On CI, `settings-general` read 4 tokens twice and then
|
|
313
|
+
// scored 1.00 against the *Settings root* — because the General screen's
|
|
314
|
+
// back button is labelled "Settings", so a General seen only down to its
|
|
315
|
+
// nav bar is structurally the same screen as the root. The harness
|
|
316
|
+
// reported it as a fingerprint that does not resemble its own screen. The
|
|
317
|
+
// fingerprint was fine; the look was too short. `settings-general` reads
|
|
318
|
+
// 5 tokens when it is read properly.
|
|
319
|
+
//
|
|
320
|
+
// Best-of, not last, because these reads are samples of a screen that is
|
|
321
|
+
// still drawing: a later read is usually richer but not reliably so, and
|
|
322
|
+
// throwing away a 5-token reading because the retry saw 4 would be the
|
|
323
|
+
// same bug with more steps.
|
|
324
|
+
const before = (id.tokens ?? []).length;
|
|
325
|
+
const seen = [before];
|
|
326
|
+
for (let attempt = 1; attempt <= SPARSE_READ_RETRIES; attempt += 1) {
|
|
327
|
+
await new Promise((r) => setTimeout(r, SPARSE_READ_WAIT_MS * attempt));
|
|
328
|
+
const again = await api.screenIdentity(device, { fresh: true, confirmNovel: false });
|
|
329
|
+
seen.push((again.tokens ?? []).length);
|
|
330
|
+
if ((again.tokens ?? []).length > (id.tokens ?? []).length) id = again;
|
|
331
|
+
if ((id.tokens ?? []).length >= MIN_TOKENS_FOR_A_READING) break;
|
|
332
|
+
}
|
|
333
|
+
const after = (id.tokens ?? []).length;
|
|
334
|
+
console.log(` ("${screen.name}" read ${before} token(s) — too sparse to compare;`
|
|
335
|
+
+ ` read again: ${seen.slice(1).join(', ')} — kept ${after})`);
|
|
336
|
+
if (after < MIN_TOKENS_FOR_A_READING) {
|
|
337
|
+
sparseReadings.push(`${screen.name} round ${round}: ${after} token(s) after ${seen.length} reads`);
|
|
338
|
+
sparseAt.add(`${screen.name}|${round}`);
|
|
339
|
+
}
|
|
340
|
+
}
|
|
152
341
|
// Did we actually arrive? Two differently-named screens reading the same
|
|
153
342
|
// fingerprint means the navigation did not land before the reading was
|
|
154
343
|
// taken, and every distribution below it is then measuring the tour rather
|
|
@@ -219,9 +408,45 @@ for (let round = 1; round <= rounds; round += 1) {
|
|
|
219
408
|
}
|
|
220
409
|
}
|
|
221
410
|
|
|
222
|
-
save();
|
|
411
|
+
save({ staleReadings });
|
|
223
412
|
if (outFile) console.log(`\nwrote ${readings.length} readings to ${outFile}`);
|
|
224
413
|
|
|
414
|
+
if (sparseReadings.length) {
|
|
415
|
+
// Reported, NOT failed — and that correction is worth more than the check.
|
|
416
|
+
//
|
|
417
|
+
// This exited 1 for four hours on 2026-09-14 and turned an intermittent CI
|
|
418
|
+
// failure into a deterministic one. The reasoning was that a 4-token reading
|
|
419
|
+
// is a screen that has not drawn. It is not: the Settings root legitimately
|
|
420
|
+
// reads 4 tokens on a hosted runner, twice in a row, three times in a row —
|
|
421
|
+
// which the comment on `MIN_TOKENS_FOR_A_READING` had itself said ("4-8
|
|
422
|
+
// tokens") one screen above the code that rejected it.
|
|
423
|
+
//
|
|
424
|
+
// What remains true is that a reading this bare cannot distinguish its screen
|
|
425
|
+
// from anything else equally bare, so it is worth saying out loud. What is not
|
|
426
|
+
// true is that saying so should stop the run. The arrival check below is the
|
|
427
|
+
// one that catches the real collision, and it does so on evidence rather than
|
|
428
|
+
// on a token count.
|
|
429
|
+
console.log(`\nNOTE ${sparseReadings.length} reading(s) were too bare to distinguish a screen:`);
|
|
430
|
+
for (const f of sparseReadings) console.log(` ${f}`);
|
|
431
|
+
console.log(`\nFewer than ${MIN_TOKENS_FOR_A_READING} tokens after two reads. That is not automatically`);
|
|
432
|
+
console.log('wrong — a plain screen really can be this bare — but two readings this sparse');
|
|
433
|
+
console.log('cannot be told apart, which is the hazard TOKEN_RULES_VERSION 7 was written');
|
|
434
|
+
console.log('for. If a same-screen score below disagrees with itself, start here.');
|
|
435
|
+
}
|
|
436
|
+
|
|
437
|
+
if (staleReadings.length) {
|
|
438
|
+
// A NOTE and not a failure, for the same reason the sparseness guard is one:
|
|
439
|
+
// this says the runner was too slow to give a fresh frame, which is a fact
|
|
440
|
+
// about the machine and not about the fingerprint. It is printed *above* the
|
|
441
|
+
// arrival check on purpose — when both fire, this is the explanation of that.
|
|
442
|
+
console.log(`\nNOTE ${staleReadings.length} reading(s) never got a frame newer than their navigation:`);
|
|
443
|
+
for (const f of staleReadings) console.log(` ${f}`);
|
|
444
|
+
console.log(`\nCapture on this host is behind the tour. With damage-driven capture a "fresh"`);
|
|
445
|
+
console.log('read returns the newest frame that exists, so on a slow runner it can predate');
|
|
446
|
+
console.log('the navigation entirely. If an arrival failure follows, this is its cause and');
|
|
447
|
+
console.log('the tour is not what needs fixing.');
|
|
448
|
+
}
|
|
449
|
+
|
|
225
450
|
if (arrivalFailures.length) {
|
|
226
451
|
console.error(`\nFAIL ${arrivalFailures.length} reading(s) were taken on the previous screen:`);
|
|
227
452
|
for (const f of arrivalFailures) console.error(` ${f}`);
|
|
@@ -296,6 +521,20 @@ if (strays.length) {
|
|
|
296
521
|
const matchNamed = match ? namedTokens(match.tokens) : [];
|
|
297
522
|
const collided = bestOther >= 0.99 && named.length === 0 && matchNamed.length === 0;
|
|
298
523
|
if (collided) collisions += 1;
|
|
524
|
+
// The third cause, and the one that produced this report on 2026-09-15.
|
|
525
|
+
//
|
|
526
|
+
// A reading that stayed under the token floor did not fail to resemble its
|
|
527
|
+
// screen — it never saw enough of the screen to resemble anything. On CI
|
|
528
|
+
// `settings-general` read 4 tokens and scored 1.00 against the Settings
|
|
529
|
+
// root, because the General screen's back button is labelled "Settings", so
|
|
530
|
+
// a General seen only down to its nav bar IS the root structurally. The
|
|
531
|
+
// report called that a wrong turn and sent the reader to fix the tour.
|
|
532
|
+
//
|
|
533
|
+
// Not folded into `collided` above: that one means the fingerprint had
|
|
534
|
+
// nothing to work with, which is this harness's subject. This means we did
|
|
535
|
+
// not look long enough, which is the harness's own fault and a different
|
|
536
|
+
// remedy.
|
|
537
|
+
const underRead = sparseAt.has(`${reading.name}|${reading.round}`);
|
|
299
538
|
console.error(` ${reading.name} r${reading.round}: own screen ${bestSelf.toFixed(2)}, `
|
|
300
539
|
+ `${match ? `${match.name} r${match.round}` : 'another screen'} ${bestOther.toFixed(2)} `
|
|
301
540
|
+ `(${reading.count} tokens, ${named.length} named, sources ${reading.sources.join('+') || 'none'})`);
|
|
@@ -303,6 +542,10 @@ if (strays.length) {
|
|
|
303
542
|
console.error(' ^ a COLLISION, not a wrong turn: neither reading carries a chrome');
|
|
304
543
|
console.error(' label, so both are structure with no name and the fingerprint has');
|
|
305
544
|
console.error(' nothing left to tell two list screens apart.');
|
|
545
|
+
} else if (underRead) {
|
|
546
|
+
console.error(' ^ UNDER-READ, not a wrong turn: this reading stayed below the token');
|
|
547
|
+
console.error(' floor after every retry, so it never saw enough of its screen to');
|
|
548
|
+
console.error(' resemble one. Two screens read this thinly are the same screen.');
|
|
306
549
|
}
|
|
307
550
|
if (named.length) console.error(` names: ${named.map((t) => t.slice(t.indexOf('"'), t.lastIndexOf('"') + 1)).join(' ')}`);
|
|
308
551
|
}
|
|
@@ -310,6 +553,10 @@ if (strays.length) {
|
|
|
310
553
|
console.error(`\n${collisions} of ${strays.length} are fingerprint collisions. That is this harness's own subject,`);
|
|
311
554
|
console.error('not a tour fault: a reading whose chrome label went missing cannot establish');
|
|
312
555
|
console.error('identity, and comparing it as though it could is what produced the verdict above.');
|
|
556
|
+
} else if (strays.every((x) => sparseAt.has(`${x.reading.name}|${x.reading.round}`))) {
|
|
557
|
+
console.error('\nEvery stray above was UNDER-READ, so this says nothing about the tour or the');
|
|
558
|
+
console.error('fingerprint — the harness scored a look that was too short. The retries are in');
|
|
559
|
+
console.error('SPARSE_READ_RETRIES; a runner that needs more than they allow is the finding.');
|
|
313
560
|
} else {
|
|
314
561
|
console.error('\nThat is the tour going somewhere unintended, not the fingerprint drifting, and');
|
|
315
562
|
console.error('measuring it as either distribution poisons both ends. Fix the tour — a tap that');
|