simframe 0.16.0 → 0.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +90 -11
- package/package.json +1 -1
- package/scripts/bench-hpi.mjs +8 -0
- package/scripts/ci-device-guard.mjs +27 -0
- package/scripts/ci-memory.mjs +24 -2
- package/scripts/device-state.mjs +5 -62
- package/src/actions.js +118 -3
- package/src/cli.js +84 -4
- package/src/device-state.js +64 -0
- package/src/graph.js +1 -1
- package/src/matching.js +3 -2
- package/src/mcp.js +24 -2
- package/src/metrics.js +144 -3
- package/src/navigate.js +59 -3
- package/src/wedge.js +218 -0
package/README.md
CHANGED
|
@@ -172,6 +172,18 @@ selector that did not *resolve* and nothing else, because retrying from a screen
|
|
|
172
172
|
you did not expect to be on is a second guess. A destructive-looking label is
|
|
173
173
|
refused as a substitute even if you list it.
|
|
174
174
|
|
|
175
|
+
**Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
|
|
176
|
+
"optional": true}` is skipped when nothing matches and runs normally when
|
|
177
|
+
something does. It exists because the opposite cost real time: a batch that
|
|
178
|
+
included a dismissal for a first-launch sheet lost **six correct steps** on the
|
|
179
|
+
next run because the sheet *did not appear*. A nag screen, a permission prompt,
|
|
180
|
+
a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
|
|
181
|
+
is the expensive outcome — a call per step instead of a call per flow. Only a
|
|
182
|
+
selector that resolved to **nothing** is absorbed: a target that is on screen
|
|
183
|
+
twice is ambiguous, not absent, and must still verify. Skipped steps are
|
|
184
|
+
reported, because "the sheet was gone" and "the sheet was dismissed" are
|
|
185
|
+
different facts.
|
|
186
|
+
|
|
175
187
|
**`{"seek": "change username", "budget": 6}`** opens containers, checks, and
|
|
176
188
|
comes back, depth first, inside a hard budget. It **acts** — opening a door
|
|
177
189
|
changes state — and it refuses to open anything that commits, abandons or
|
|
@@ -400,7 +412,7 @@ steer the model is a tool surface the model uses wrong.
|
|
|
400
412
|
| `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
|
|
401
413
|
| `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
|
|
402
414
|
| `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
|
|
403
|
-
| `sim_flow_run` | Replay a flow
|
|
415
|
+
| `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
|
|
404
416
|
| `sim_find` | Resolve an intent to one control, without acting on it. |
|
|
405
417
|
| `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
|
|
406
418
|
| `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
|
|
@@ -710,16 +722,59 @@ feels, and it took two field reports to see it. Per-step wall clock is
|
|
|
710
722
|
(model round trip + simframe work × n) / n for n steps in one call
|
|
711
723
|
```
|
|
712
724
|
|
|
713
|
-
|
|
714
|
-
|
|
715
|
-
|
|
716
|
-
|
|
717
|
-
|
|
|
718
|
-
|
|
|
719
|
-
| a
|
|
720
|
-
|
|
|
721
|
-
|
|
722
|
-
|
|
725
|
+
**Compare like with like, which this table used not to.** A human tester's
|
|
726
|
+
1.95 s per step is the *whole* loop — look at the screen, decide what to do,
|
|
727
|
+
do it. Any row that does not include a decision is not comparable to it.
|
|
728
|
+
|
|
729
|
+
| | perceives | decides | acts | per step |
|
|
730
|
+
| --- | --- | --- | --- | --- |
|
|
731
|
+
| **a human tester, measured** | yes | yes | yes | **1.95 s** |
|
|
732
|
+
| **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
|
|
733
|
+
| simframe, batch of 4 | yes | yes | yes | ~6.7 s |
|
|
734
|
+
| simframe, one model call per step | yes | yes | yes | ~21.7 s |
|
|
735
|
+
| — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
|
|
736
|
+
| — a saved flow replayed | yes | **no** | yes | 1.98 s |
|
|
737
|
+
|
|
738
|
+
**The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
|
|
739
|
+
a human**, and about eleven times when a failure forces one model call per step.
|
|
740
|
+
|
|
741
|
+
The two indented rows are the ones this README used to lead with, and both are
|
|
742
|
+
category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
|
|
743
|
+
taken out — the thinking is the model round trip, which is most of the clock.
|
|
744
|
+
And a replay *decides nothing*: it is a recording being played back, so its fair
|
|
745
|
+
counterpart is a human repeating a flow they have memorised, who would be well
|
|
746
|
+
under 1.95 s. Replay against a human working something out for the first time is
|
|
747
|
+
a rehearsal measured against a first attempt.
|
|
748
|
+
|
|
749
|
+
What the two rows do establish, and it is the finding that reordered this
|
|
750
|
+
project: **there is nothing left to win inside the engine.** 1.7 s is small
|
|
751
|
+
beside a 20 s round trip, so making perception or input faster buys single-digit
|
|
752
|
+
percentages. The only variable that matters is `n` — how many steps one decision
|
|
753
|
+
covers. Every improvement here has come from raising it, not from faster code.
|
|
754
|
+
|
|
755
|
+
### One more term: the launch is a fixed cost
|
|
756
|
+
|
|
757
|
+
Measured on the benchmark suite with **no model in the loop at all**:
|
|
758
|
+
|
|
759
|
+
| route | steps | agent | per step | human | per step |
|
|
760
|
+
| --- | --- | --- | --- | --- | --- |
|
|
761
|
+
| contacts-kate-bell | 2 | 12419 ms | **6.2 s** | 4300 ms | 2.15 s |
|
|
762
|
+
| settings-larger-text | 4 | 15281 ms | 3.8 s | 7799 ms | 1.95 s |
|
|
763
|
+
|
|
764
|
+
2.9× a human on the short route with nothing thinking, which does not fit
|
|
765
|
+
`1.7s × n`. A cold app launch is a **one-off cost amortised over the route**:
|
|
766
|
+
|
|
767
|
+
```
|
|
768
|
+
per step = (model round trip + launch cost + ~1.7s × n) / n
|
|
769
|
+
```
|
|
770
|
+
|
|
771
|
+
That reconciles a 7-step replay at 1.82 s/step with a 2-step route at 6.2 s/step
|
|
772
|
+
— one launch spread over 7 steps or over 2. **The `~1.7 s` figure is warm taps
|
|
773
|
+
inside a batch**, and short routes are materially worse than the table above
|
|
774
|
+
implies on its own.
|
|
775
|
+
|
|
776
|
+
`step_ratio` was **1** throughout: when a run completed it took exactly the
|
|
777
|
+
minimum number of steps.
|
|
723
778
|
A field report put the split at **34% simframe, 60% agent round trips** over
|
|
724
779
|
462 s of wall clock — the tester's *"30+ seconds between each step"* was
|
|
725
780
|
accurate and was not simframe. So there is nothing left to win inside the
|
|
@@ -840,6 +895,7 @@ simframe supervisions # local supervisor rulings, and what came of each
|
|
|
840
895
|
simframe hpi # speed and accuracy against a human baseline
|
|
841
896
|
simframe baseline record settings-larger-text --runs=5 # record the human
|
|
842
897
|
simframe input reset # rebuild the HID session, without restarting anything
|
|
898
|
+
simframe diagnose # what this device is doing right now, and which failure it is
|
|
843
899
|
simframe revive # power-cycle a wedged device: stop, shutdown, boot, start, reset input
|
|
844
900
|
simframe start / status / stop [--force] / devices
|
|
845
901
|
simframe ui --device=emulator-5554 # or export SIMFRAME_DEVICE once
|
|
@@ -864,6 +920,29 @@ rebinds, no frame since). This needs the device restarted —
|
|
|
864
920
|
`simframe revive --device=<udid>`. Backing off until a frame arrives.
|
|
865
921
|
```
|
|
866
922
|
|
|
923
|
+
**`simframe diagnose` says which failure it is**, which `doctor` cannot: `doctor`
|
|
924
|
+
answers "can this machine capture", and this answers "what is this device doing
|
|
925
|
+
right now". It reads only what discriminates — frame sequence, age and
|
|
926
|
+
stillness, element counts split by sensor, how many of them fuse, and who holds
|
|
927
|
+
the front by pid — and returns one of:
|
|
928
|
+
|
|
929
|
+
| verdict | what it means |
|
|
930
|
+
| --- | --- |
|
|
931
|
+
| `capture-down` | no frames at all, carrying the daemon's own sentence |
|
|
932
|
+
| `nothing-readable` | frames arriving, neither sensor finds a single element |
|
|
933
|
+
| `stale-frame` | both sensors full, almost nothing fuses — the framebuffer is behind the tree, so **an image from this device is not safe to trust** |
|
|
934
|
+
| `not-presenting` | an app holds the front by pid and the display shows almost nothing |
|
|
935
|
+
| `healthy` | — |
|
|
936
|
+
|
|
937
|
+
`not-presenting` deliberately does **not** say whether that is a lock screen, a
|
|
938
|
+
dead surface or a crashed system shell. It is not knowable from here, and
|
|
939
|
+
guessing is how a regex ended up standing where a measurement belongs.
|
|
940
|
+
|
|
941
|
+
The `stale-frame` threshold is measured rather than chosen: element fusion on
|
|
942
|
+
five healthy screens ran 0.667–0.929, so the threshold sits at 0.1 — 6.7× below
|
|
943
|
+
the observed floor rather than inside the metric's own noise. Numbers in
|
|
944
|
+
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
945
|
+
|
|
867
946
|
`simframe revive` is that restart, in the order that matters — stop the daemon,
|
|
868
947
|
shut the device down, boot it and *wait for the boot to finish*, start capture,
|
|
869
948
|
rebuild the HID session — and it ends by checking frames are flowing again
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.18.0",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
5
|
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
package/scripts/bench-hpi.mjs
CHANGED
|
@@ -191,6 +191,14 @@ for (const f of report.flows) {
|
|
|
191
191
|
report.overall.hpi_time_median_of_passes = passTimes.length ? Number(metrics.median(passTimes).toFixed(3)) : null;
|
|
192
192
|
const o = report.overall;
|
|
193
193
|
console.log(`\nHPI_accuracy ${o.hpi_accuracy} HPI_time ${o.hpi_time ?? '—'} HPI ${o.hpi ?? '—'} step_ratio ${o.step_ratio ?? '—'}`);
|
|
194
|
+
if (o.runs_lost_to_device) {
|
|
195
|
+
// Measured on a full local suite: 5 of 17 runs failed because the guest's
|
|
196
|
+
// SpringBoard crashed or simctl stopped answering. They used to lower
|
|
197
|
+
// HPI_accuracy, which is the number this job gates on.
|
|
198
|
+
console.log(` over ${o.runs} measurable run(s); ${o.runs_lost_to_device} left the denominator`
|
|
199
|
+
+ ' because the DEVICE failed, not the code (item 173):');
|
|
200
|
+
for (const c of o.device_causes ?? []) console.log(` ${c}`);
|
|
201
|
+
}
|
|
194
202
|
if (passTimes.length > 1) {
|
|
195
203
|
console.log(`HPI_time per pass: ${passTimes.join(', ')} — median ${o.hpi_time_median_of_passes} (what the gate reads)`);
|
|
196
204
|
}
|
|
@@ -59,6 +59,25 @@ if (!cause) {
|
|
|
59
59
|
process.exit(first.code ?? 1);
|
|
60
60
|
}
|
|
61
61
|
|
|
62
|
+
// Diagnose BEFORE reviving, because reviving is what destroys the evidence.
|
|
63
|
+
//
|
|
64
|
+
// This guard has been recognising wedges from the shape of the failure text
|
|
65
|
+
// since 126, and then immediately power-cycling the device — so item 173 has
|
|
66
|
+
// accumulated a dozen occurrences and not one observation of what the device
|
|
67
|
+
// was doing at the time. `cause` above names a *consequence* ("a launched app
|
|
68
|
+
// never came to the front"); this names what the two sensors actually saw, and
|
|
69
|
+
// the difference decides between a stale framebuffer, a screen that is not the
|
|
70
|
+
// app, and capture being down. Its cost is one read on a path that is already
|
|
71
|
+
// failing.
|
|
72
|
+
const diagnose = async (when) => {
|
|
73
|
+
const d = await run(['node', 'src/cli.js', 'diagnose', `--device=${udid}`]);
|
|
74
|
+
console.error(`\n (diagnosis ${when} — for DEFERRED 173)`);
|
|
75
|
+
const verdict = /— (\S+)\n/.exec(d.out)?.[1] ?? 'unreadable';
|
|
76
|
+
summary(` - diagnosis ${when}: \`${verdict}\``);
|
|
77
|
+
return verdict;
|
|
78
|
+
};
|
|
79
|
+
const before = await diagnose('before the revive');
|
|
80
|
+
|
|
62
81
|
console.error(`\n (${cause} — DEFERRED 126. Reviving once and running again.)`);
|
|
63
82
|
await run(['node', 'src/cli.js', 'revive', `--device=${udid}`], { capture: false });
|
|
64
83
|
const second = await run(cmd);
|
|
@@ -70,4 +89,12 @@ if (second.code === 0) {
|
|
|
70
89
|
const again = deviceCause(second.out);
|
|
71
90
|
summary(`- \`${cmd.join(' ')}\` — **${again ? 'device unavailable' : 'check failed'}** after a revive${again ? ` (${again})` : ''}`);
|
|
72
91
|
if (again) console.error(`\nFAIL the simulator is still in a bad state after a revive: ${again}`);
|
|
92
|
+
// Twice is the interesting case: a revive cured it and it came back, or the
|
|
93
|
+
// revive did not cure it at all. Those are different faults and the pair of
|
|
94
|
+
// diagnoses says which.
|
|
95
|
+
const after = await diagnose('after the revive');
|
|
96
|
+
if (before !== after) {
|
|
97
|
+
console.error(`\n (the device changed state across the revive: ${before} -> ${after})`);
|
|
98
|
+
summary(` - state changed across the revive: \`${before}\` -> \`${after}\``);
|
|
99
|
+
}
|
|
73
100
|
process.exit(second.code ?? 1);
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -646,12 +646,25 @@ console.log('\n--- what may be saved, and what may not ---');
|
|
|
646
646
|
// evidence for it. Checked against whatever the run actually produced rather
|
|
647
647
|
// than assuming a shape, as before.
|
|
648
648
|
const attempt = await jsonRetry(['do', LOOP, `--save=${FLOW_NAME}`], { allowFail: true });
|
|
649
|
+
//
|
|
650
|
+
// Three outcomes, not two. The missing one is a step that ran and *failed*: a
|
|
651
|
+
// failing step stops the batch, so a failure on the last step leaves
|
|
652
|
+
// `ranSteps === steps.length` and the run reads as complete. A peer watched
|
|
653
|
+
// `FLOW FAILED — 4 ok, 1 failed (of 5)` save itself. This branch is also the
|
|
654
|
+
// one that goes stale: yesterday the save contract was replaced, the unit test
|
|
655
|
+
// was updated, and this file was left asserting the old rule — so the three
|
|
656
|
+
// cases are enumerated here explicitly rather than left as an `else`.
|
|
649
657
|
const contradicted = attempt.results.some((r) => /^unexpected/.test(r.verification?.verdict ?? ''));
|
|
658
|
+
const anyFailed = attempt.results.some((r) => !r.ok);
|
|
650
659
|
const clean = attempt.results.every((r) => !r.verification || r.verification.verdict === 'ok');
|
|
651
660
|
if (contradicted) {
|
|
652
661
|
check(attempt.saved?.ok === false && attempt.saved?.reason === 'contradicted-steps',
|
|
653
662
|
'a flow with a contradicted step is refused, not quietly saved',
|
|
654
663
|
`${attempt.saved?.reason} (${(attempt.saved?.verdicts ?? []).join(', ')})`);
|
|
664
|
+
} else if (anyFailed) {
|
|
665
|
+
check(attempt.saved?.ok === false && attempt.saved?.reason === 'failed-steps',
|
|
666
|
+
'a flow with a step that ran and failed is refused, not quietly saved',
|
|
667
|
+
`${attempt.saved?.reason} (failed step ${(attempt.saved?.failed ?? []).join(', ')})`);
|
|
655
668
|
} else {
|
|
656
669
|
check(attempt.saved?.ok === true, 'a first traversal is recordable',
|
|
657
670
|
`${attempt.saved?.steps} steps`);
|
|
@@ -676,12 +689,21 @@ if (check(forced.saved?.ok === true, 'and --force saves it anyway', `${forced.sa
|
|
|
676
689
|
// state nothing ever leaves: a clean replay is the confirmation a first
|
|
677
690
|
// traversal could not give. Only asserted when the replay actually ran to
|
|
678
691
|
// the end — a partial replay promotes nothing, deliberately.
|
|
692
|
+
const after = await jsonRetry(['flow', 'list']);
|
|
693
|
+
const entry = after.find((f) => f.name === FLOW_NAME);
|
|
679
694
|
if (replayed.ok) {
|
|
680
|
-
const after = await jsonRetry(['flow', 'list']);
|
|
681
|
-
const entry = after.find((f) => f.name === FLOW_NAME);
|
|
682
695
|
check(entry && !entry.provisional,
|
|
683
696
|
'and a clean replay confirms a provisional flow',
|
|
684
697
|
`provisional=${entry?.provisional ?? 'gone'}`);
|
|
698
|
+
} else {
|
|
699
|
+
// The other direction, and it is the one that was silently wrong: promotion
|
|
700
|
+
// keyed on `ranSteps === steps.length`, which a failure on the *last* step
|
|
701
|
+
// satisfies. Both of a peer's saved flows were marked confirmed by replays
|
|
702
|
+
// that failed. A confirmation that a failing replay can grant is not a
|
|
703
|
+
// confirmation, so the negative case has to be checked too.
|
|
704
|
+
check(entry?.provisional === true,
|
|
705
|
+
'and a replay that failed does NOT confirm the flow it just disproved',
|
|
706
|
+
`ok=${replayed.ok}, provisional=${entry?.provisional}`);
|
|
685
707
|
}
|
|
686
708
|
const unknown = await cli(['flow', 'run', 'no-such-flow'], { expectFail: true });
|
|
687
709
|
check(/no flow/i.test(unknown), 'an unknown flow name is refused with what is known');
|
package/scripts/device-state.mjs
CHANGED
|
@@ -1,62 +1,5 @@
|
|
|
1
|
-
//
|
|
2
|
-
//
|
|
3
|
-
//
|
|
4
|
-
//
|
|
5
|
-
|
|
6
|
-
// and the only thing that exercised it was a hosted runner, at the end of a
|
|
7
|
-
// fifteen-minute job, in the middle of a report. Two runtime bugs in this
|
|
8
|
-
// project came from logic that was correct and had never executed.
|
|
9
|
-
//
|
|
10
|
-
// Every entry here has been seen in a real run. Adding one from imagination is
|
|
11
|
-
// how a guard starts reviving genuine failures into passes.
|
|
12
|
-
|
|
13
|
-
/** Conditions that are the simulator, not the code. Each seen in a real run. */
|
|
14
|
-
export const DEVICE_STATE = [
|
|
15
|
-
[/NSPOSIXErrorDomain.*code=?\s*60|Operation timed out/i, 'simctl stopped answering (NSPOSIXErrorDomain 60)'],
|
|
16
|
-
[/did not produce a frame|produced no frame in \d+s/i, 'the daemon is up and the display renders nothing'],
|
|
17
|
-
[/Timeout waiting for screen surfaces|display surface is not answering|display surface could not be read/i, 'the display surface is wedged'],
|
|
18
|
-
[/no frames buffered|capture is wedged/i, 'capture stopped'],
|
|
19
|
-
[/the second app never launched|could not be dispatched/i, 'an app would not launch'],
|
|
20
|
-
// A launched app that never comes to the front, seen as the tour waiting for
|
|
21
|
-
// one of its landmarks on a screen that is showing a clock and nothing else.
|
|
22
|
-
//
|
|
23
|
-
// Measured on a runner: `ok launch — launched com.apple.Preferences
|
|
24
|
-
// (relaunched)` followed by `waited 8000ms for General: "General" is not on
|
|
25
|
-
// this screen. Visible: 10:50, .?o (the screen has not moved for 6181ms)`.
|
|
26
|
-
// Two labels, one of them a clock, on a still screen — the device is not
|
|
27
|
-
// presenting the app, and the guard called that a check failing on its
|
|
28
|
-
// merits and declined to revive.
|
|
29
|
-
//
|
|
30
|
-
// Deliberately narrow. It requires the wait to have failed AND the screen to
|
|
31
|
-
// have been still AND almost nothing readable: a tour that genuinely asks for
|
|
32
|
-
// the wrong label has a screen full of other labels, and must keep failing
|
|
33
|
-
// rather than being retried into a pass.
|
|
34
|
-
[
|
|
35
|
-
/never arrived[\s\S]*?Visible:[^\n]{0,24}\(the screen has not moved for \d+ms/i,
|
|
36
|
-
'a launched app never came to the front (the screen shows a clock and nothing else)',
|
|
37
|
-
],
|
|
38
|
-
// The same condition, now said outright by the step that suffered it instead
|
|
39
|
-
// of inferred from the shape of the screen afterwards. Item 169 gave `launch`
|
|
40
|
-
// a pid to compare, so a launch that starts a process and never fronts it
|
|
41
|
-
// reports itself; this signature fires on the cause rather than on a
|
|
42
|
-
// consequence that had to be recognised by "two labels, one a clock".
|
|
43
|
-
//
|
|
44
|
-
// It cannot be triggered by a tour asking for the wrong label — only a failed
|
|
45
|
-
// launch emits this sentence — so it needs none of the narrowing above.
|
|
46
|
-
[
|
|
47
|
-
/never came to the front within \d+ms/i,
|
|
48
|
-
'a launched app never came to the front (the launch said so itself, by pid)',
|
|
49
|
-
],
|
|
50
|
-
// Seen on the v0.14.3 bench run: `could not launch com.apple.Preferences:
|
|
51
|
-
// The system shell (SpringBoard:36454) probably crashed.` The guest's window
|
|
52
|
-
// server going down is the device, not the check, and nothing here matched it.
|
|
53
|
-
[
|
|
54
|
-
/system shell \(SpringBoard[^)]*\) probably crashed/i,
|
|
55
|
-
"the guest's SpringBoard crashed, so nothing can be fronted",
|
|
56
|
-
],
|
|
57
|
-
];
|
|
58
|
-
|
|
59
|
-
/** The condition this output shows, or null when the check failed on its merits. */
|
|
60
|
-
export function deviceCause(text) {
|
|
61
|
-
return DEVICE_STATE.find(([re]) => re.test(String(text ?? '')))?.[1] ?? null;
|
|
62
|
-
}
|
|
1
|
+
// Kept as a re-export: the table moved to `src/device-state.js` so `src/` can
|
|
2
|
+
// use it without importing out of `scripts/`. `ci-device-guard.mjs` and the
|
|
3
|
+
// unit test both reach it through this path, and a redirect is cheaper than
|
|
4
|
+
// updating every caller for a move that changes nothing about the table.
|
|
5
|
+
export { DEVICE_STATE, deviceCause } from '../src/device-state.js';
|
package/src/actions.js
CHANGED
|
@@ -346,6 +346,16 @@ export async function runScript(
|
|
|
346
346
|
// the caller returns to Claude is rendered from this, so describing the end
|
|
347
347
|
// state costs nothing beyond the verification pass the flow already ran.
|
|
348
348
|
let endScreen = null;
|
|
349
|
+
// Where this run began, so a flow saved from it can say what it assumes.
|
|
350
|
+
//
|
|
351
|
+
// Null when the first step does not depend on where we were — see
|
|
352
|
+
// `resetsTheScreen`.
|
|
353
|
+
//
|
|
354
|
+
// It was recorded as `null` on every flow 0.16.0 saved, which a peer caught:
|
|
355
|
+
// "replay never checks it is starting where it was recorded". Reading it is
|
|
356
|
+
// free — the first step computes the identity anyway — and it is a fact the
|
|
357
|
+
// recording had and discarded.
|
|
358
|
+
let startScreen = null;
|
|
349
359
|
// At most one recovery per run. Pressing home while already on the springboard
|
|
350
360
|
// moves nothing and is not a failure, so an unbounded retry would rebuild the
|
|
351
361
|
// session and press again on every such step for no reason.
|
|
@@ -393,6 +403,9 @@ export async function runScript(
|
|
|
393
403
|
? (carriedScreen ?? await api.screenIdentity(deviceQuery, { options, settleMs: stableMs, timeoutMs, confirmNovel }))
|
|
394
404
|
: null;
|
|
395
405
|
carriedScreen = null;
|
|
406
|
+
if (i === 0 && !resetsTheScreen(step)) {
|
|
407
|
+
startScreen = beforeScreen ? { hash: beforeScreen.hash ?? null, tokens: beforeScreen.tokens ?? null } : null;
|
|
408
|
+
}
|
|
396
409
|
// What this action did last time it was taken here, if ever.
|
|
397
410
|
const prediction = verify && beforeScreen?.hash ? graph.predict(udid, beforeScreen, step) : null;
|
|
398
411
|
try {
|
|
@@ -920,7 +933,19 @@ export async function runScript(
|
|
|
920
933
|
// switch moving 0.1% of the screen, which is neither faculty. Saying
|
|
921
934
|
// "assumed" is the honest answer; guessing a better-sounding reason
|
|
922
935
|
// would be the same mistake in the other direction.
|
|
923
|
-
|
|
936
|
+
//
|
|
937
|
+
// **Per verdict, though, not per site.** That argument is about
|
|
938
|
+
// `no-visible-change` and does not extend to every verdict: an
|
|
939
|
+
// `unexpected-screen` means the screen after the action was not the
|
|
940
|
+
// one memory predicted, which is item 174 and nothing else.
|
|
941
|
+
// `metrics.VERDICT_FACULTY` holds the verdicts whose faculty is read,
|
|
942
|
+
// and `no-visible-change` is deliberately not one of them — so this
|
|
943
|
+
// still says "assumed" for the 162 records the reasoning above is
|
|
944
|
+
// actually about, and stops saying it for the 26 it never covered.
|
|
945
|
+
classified: Boolean(metrics.VERDICT_FACULTY[verification?.verdict]),
|
|
946
|
+
// Recorded as a field rather than left as a prefix of `detail`, which
|
|
947
|
+
// is how the breakdown had to recover it: by parsing a string.
|
|
948
|
+
verdict: verification?.verdict ?? null,
|
|
924
949
|
// `verification_failed` is the largest reason class in the log and it
|
|
925
950
|
// was the only one carrying no intent, which made most of the corpus
|
|
926
951
|
// useless for asking what kind of decision costs us. The step knows
|
|
@@ -945,10 +970,42 @@ export async function runScript(
|
|
|
945
970
|
break;
|
|
946
971
|
}
|
|
947
972
|
} catch (err) {
|
|
973
|
+
// A step that was allowed not to be there, and is not there.
|
|
974
|
+
//
|
|
975
|
+
// `or:` supplies alternative *selectors*; nothing could express "skip
|
|
976
|
+
// this if absent". So a batch crossing a first-launch nag, a permission
|
|
977
|
+
// prompt or a "What's New" could not be batched at all — the interstitial
|
|
978
|
+
// appears on run one and not on run two, the step fails, and every
|
|
979
|
+
// remaining step is discarded. A peer lost six correct steps because a
|
|
980
|
+
// sheet *did not* appear, and on the framing that every hard-fail
|
|
981
|
+
// dropping a caller back to single-stepping is a latency bug, that is the
|
|
982
|
+
// most expensive shape there is: ~20 s per step, for the rest of the plan.
|
|
983
|
+
//
|
|
984
|
+
// Deliberately narrow. Only `unknown_screen` — the selector resolved to
|
|
985
|
+
// nothing — absorbs. `ambiguous_intent` means the target *is* there,
|
|
986
|
+
// twice, and a step that is present must still verify; so must one that
|
|
987
|
+
// resolved and then failed. "Skip if absent" must not become "tap
|
|
988
|
+
// whatever is there".
|
|
989
|
+
if (step.optional && didNotResolve(err)) {
|
|
990
|
+
results.push({
|
|
991
|
+
index: i,
|
|
992
|
+
action: step.action,
|
|
993
|
+
ok: true,
|
|
994
|
+
skipped: true,
|
|
995
|
+
ms: Date.now() - stepStart,
|
|
996
|
+
detail: `skipped — optional, and ${JSON.stringify(String(goalOf(step) ?? step.action))} is not on this screen`,
|
|
997
|
+
});
|
|
998
|
+
continue;
|
|
999
|
+
}
|
|
948
1000
|
results.push({ index: i, action: step.action, ok: false, ms: Date.now() - stepStart, error: err.message });
|
|
949
1001
|
const why = metrics.reasonForStepError(step, err);
|
|
950
1002
|
noteEscalation({
|
|
951
1003
|
stepIndex: i,
|
|
1004
|
+
// When the failure is the simulator rather than the code, say so. 78 of
|
|
1005
|
+
// the bench device's `verification_failed` records are `simctl` failing,
|
|
1006
|
+
// an app that would not launch, or capture stopping — item 173, counted
|
|
1007
|
+
// towards a perception phase.
|
|
1008
|
+
device: why.device ?? null,
|
|
952
1009
|
fingerprint: beforeScreen?.hash ?? metrics.fingerprintNow(udid, screenmap),
|
|
953
1010
|
reason: why.reason,
|
|
954
1011
|
candidates: why.candidates,
|
|
@@ -1020,6 +1077,7 @@ export async function runScript(
|
|
|
1020
1077
|
// navigate.saveFlow without the caller reassembling what it just ran.
|
|
1021
1078
|
steps,
|
|
1022
1079
|
flowId,
|
|
1080
|
+
startScreen,
|
|
1023
1081
|
endScreen,
|
|
1024
1082
|
results,
|
|
1025
1083
|
ok: !failed,
|
|
@@ -1462,6 +1520,53 @@ export function alternativesFor(step) {
|
|
|
1462
1520
|
return list.map((v) => (typeof v === 'string' ? v : v?.value ?? v?.target ?? v?.label)).filter(Boolean);
|
|
1463
1521
|
}
|
|
1464
1522
|
|
|
1523
|
+
/**
|
|
1524
|
+
* Does this step make the screen we were on irrelevant?
|
|
1525
|
+
*
|
|
1526
|
+
* A flow that opens with `launch`, a deep link, or the home button does not
|
|
1527
|
+
* assume anything about where it starts — it goes and puts the device
|
|
1528
|
+
* somewhere. Recording a start screen for such a flow records where the device
|
|
1529
|
+
* happened to be beforehand, which is noise, and then reports a mismatch on
|
|
1530
|
+
* every replay that began anywhere else.
|
|
1531
|
+
*
|
|
1532
|
+
* Measured, not reasoned: a flow opening with `launch --relaunch` replayed
|
|
1533
|
+
* perfectly and still printed "recorded starting on 8292b488, replayed from
|
|
1534
|
+
* 39351dab", because the recording had followed a run that ended on Settings
|
|
1535
|
+
* root and the replay followed one that ended on About. Most flows open with a
|
|
1536
|
+
* launch, so that note would have fired on most correct replays — which is
|
|
1537
|
+
* item 175's defect exactly, in a feature added the same day 175 was written up.
|
|
1538
|
+
*/
|
|
1539
|
+
export function resetsTheScreen(step = {}) {
|
|
1540
|
+
if (step.action === 'launch' || step.action === 'openUrl') return true;
|
|
1541
|
+
return step.action === 'button' && /^home$/i.test(String(step.value ?? '').trim());
|
|
1542
|
+
}
|
|
1543
|
+
|
|
1544
|
+
/**
|
|
1545
|
+
* Did this failure mean "the thing is not here", as opposed to any other way a
|
|
1546
|
+
* step can fail? The one question `optional` is allowed to ask.
|
|
1547
|
+
*
|
|
1548
|
+
* **Not `reason === 'unknown_screen'`, which is what this said for one build and
|
|
1549
|
+
* what made `optional` a no-op on every screen it was built for.** The tag at
|
|
1550
|
+
* the throw site is chosen by whether the *screen* was recognised, not by
|
|
1551
|
+
* whether the *target* was found: on a screen recalled from memory an absent
|
|
1552
|
+
* target is tagged `ambiguous_intent` too. So a conditional interstitial on a
|
|
1553
|
+
* well-known screen — the whole use case — never absorbed. Caught by running it
|
|
1554
|
+
* against a device; the unit test exercised this predicate and passed, because
|
|
1555
|
+
* it tested the predicate rather than the path.
|
|
1556
|
+
*
|
|
1557
|
+
* The distinction that actually exists is `ambiguous`, set only where something
|
|
1558
|
+
* was found *several times over*, and the comment at that throw site had
|
|
1559
|
+
* already written down why it matters: "more time cannot make a thing unique,
|
|
1560
|
+
* and it can make an absent thing arrive." Absent is the case `optional` may
|
|
1561
|
+
* skip; present-twice must still verify.
|
|
1562
|
+
*/
|
|
1563
|
+
const RESOLVE_REASONS = new Set(['unknown_screen', 'ambiguous_intent']);
|
|
1564
|
+
|
|
1565
|
+
export function didNotResolve(err) {
|
|
1566
|
+
const tagged = metrics.escalationOf(err);
|
|
1567
|
+
return Boolean(tagged && RESOLVE_REASONS.has(tagged.reason) && !tagged.ambiguous);
|
|
1568
|
+
}
|
|
1569
|
+
|
|
1465
1570
|
export function mayRetryAfter(err) {
|
|
1466
1571
|
const tagged = metrics.escalationOf(err);
|
|
1467
1572
|
return Boolean(tagged && RESOLVE_FAILURES.has(tagged.reason));
|
|
@@ -3246,13 +3351,23 @@ export function settleEvidence(w) {
|
|
|
3246
3351
|
*/
|
|
3247
3352
|
export function flowSummary(res, { withTime = true } = {}) {
|
|
3248
3353
|
const time = withTime && Number.isFinite(res.totalMs) ? ` in ${res.totalMs}ms` : '';
|
|
3249
|
-
|
|
3354
|
+
// A step that was allowed not to be there and was not there did not run, and
|
|
3355
|
+
// saying so is the difference between "the nag screen was gone" and "the nag
|
|
3356
|
+
// screen was dismissed". Same lesson as the denominator above: a count that
|
|
3357
|
+
// quietly means two things is the defect.
|
|
3358
|
+
const skipped = (res.results ?? []).filter((r) => r.skipped).length;
|
|
3359
|
+
// After the time, in parentheses. Written inline before it, this read
|
|
3360
|
+
// "skipped as absent in 3412ms" — which says the skipping took 3.4 seconds.
|
|
3361
|
+
const skips = skipped
|
|
3362
|
+
? ` (${skipped} optional step${skipped === 1 ? '' : 's'} skipped as absent)`
|
|
3363
|
+
: '';
|
|
3364
|
+
if (res.ok) return `flow completed — ${res.ranSteps}/${res.totalSteps} steps${time}${skips}`;
|
|
3250
3365
|
const failed = (res.results ?? []).filter((r) => r.ok === false).length || 1;
|
|
3251
3366
|
const worked = Math.max(0, res.ranSteps - failed);
|
|
3252
3367
|
const unattempted = Math.max(0, res.totalSteps - res.ranSteps);
|
|
3253
3368
|
return `FLOW FAILED — ${worked} ok, ${failed} failed`
|
|
3254
3369
|
+ (unattempted ? `, ${unattempted} not attempted` : '')
|
|
3255
|
-
+ ` (of ${res.totalSteps})${time}`;
|
|
3370
|
+
+ ` (of ${res.totalSteps})${time}${skips}`;
|
|
3256
3371
|
}
|
|
3257
3372
|
|
|
3258
3373
|
/**
|
package/src/cli.js
CHANGED
|
@@ -11,6 +11,7 @@ import * as input from './input.js';
|
|
|
11
11
|
import * as baseline from './baseline.js';
|
|
12
12
|
import * as metrics from './metrics.js';
|
|
13
13
|
import * as navigate from './navigate.js';
|
|
14
|
+
import * as wedge from './wedge.js';
|
|
14
15
|
import { decodePng } from './png.js';
|
|
15
16
|
import * as storage from './storage.js';
|
|
16
17
|
import * as store from './store.js';
|
|
@@ -51,6 +52,7 @@ const USAGE = `simframe — always-warm iOS Simulator frames
|
|
|
51
52
|
simframe escalations [device] why simframe handed decisions back, by reason
|
|
52
53
|
simframe supervisions [device] local supervisor rulings, and what came of each
|
|
53
54
|
simframe revive [device] power-cycle a wedged device: stop, shutdown, boot, start, reset input
|
|
55
|
+
simframe diagnose [device] what this device is doing right now, and which failure it is
|
|
54
56
|
(--session=<id> narrows to one agent; the
|
|
55
57
|
ids are listed in the output. SIMFRAME_SESSION
|
|
56
58
|
names one, but only at process start — an
|
|
@@ -285,6 +287,20 @@ async function main() {
|
|
|
285
287
|
// is a different machine and should say so rather than be guessed at.
|
|
286
288
|
if (flags.readyTimeoutMs) options.readyTimeoutMs = num(flags.readyTimeoutMs);
|
|
287
289
|
|
|
290
|
+
// `--help` on any command, not only as the command.
|
|
291
|
+
//
|
|
292
|
+
// `simframe hpi --help` did the two worst things in sequence: without
|
|
293
|
+
// `--device` it demanded a device *before* printing help, and with one it
|
|
294
|
+
// ignored `--help` and ran the report. A flag that is ignored is worse than
|
|
295
|
+
// one that is rejected — it silently does something other than what was
|
|
296
|
+
// asked. (The general case, unrecognised flags being dropped everywhere
|
|
297
|
+
// rather than here, is filed separately; this is the one that was measured
|
|
298
|
+
// costing a round trip.)
|
|
299
|
+
if (flags.help || flags.h) {
|
|
300
|
+
process.stdout.write(USAGE);
|
|
301
|
+
return;
|
|
302
|
+
}
|
|
303
|
+
|
|
288
304
|
switch (command) {
|
|
289
305
|
case undefined:
|
|
290
306
|
case '-h':
|
|
@@ -433,6 +449,26 @@ async function main() {
|
|
|
433
449
|
return;
|
|
434
450
|
}
|
|
435
451
|
|
|
452
|
+
// Not folded into `doctor`, which answers "can this machine capture". This
|
|
453
|
+
// answers "what is this device doing right now", which is item 173's
|
|
454
|
+
// question and has never had an instrument.
|
|
455
|
+
case 'diagnose': {
|
|
456
|
+
const dev = await resolveDevice(device);
|
|
457
|
+
const r = await wedge.diagnose(dev.udid, { options });
|
|
458
|
+
emit(flags, r, [
|
|
459
|
+
`${r.device.name} — ${r.verdict.state}`,
|
|
460
|
+
` ${r.verdict.detail}`,
|
|
461
|
+
'',
|
|
462
|
+
` frame seq ${r.frame?.seq ?? '-'}, ${r.frame?.ageMs ?? '-'}ms old, still for ${r.frame?.stableForMs ?? '-'}ms, ${r.frame?.size ?? '-'}`,
|
|
463
|
+
` elements ${r.elements.total} total — ${r.elements.ax} by tree, ${r.elements.ocr} by OCR, ${r.elements.fused} by both`,
|
|
464
|
+
` fusion ${r.agreement ?? 'n/a'} of elements seen by both sensors (measured healthy ${wedge.HEALTHY_FUSION}; at or below ${wedge.DISAGREEMENT} the frame is stale)`,
|
|
465
|
+
` frontmost ${r.frontmost?.pid ?? 'unknown'}${r.frontmost?.title ? ` (${r.frontmost.title})` : ''}`,
|
|
466
|
+
...(r.verdict.revive ? ['', ' `simframe revive` is the recovery. Keep this output — item 173 needs it.'] : []),
|
|
467
|
+
]);
|
|
468
|
+
if (r.verdict.revive) process.exitCode = 1;
|
|
469
|
+
return;
|
|
470
|
+
}
|
|
471
|
+
|
|
436
472
|
case 'status': {
|
|
437
473
|
const udids = device
|
|
438
474
|
? [(await resolveDevice(device)).udid]
|
|
@@ -977,6 +1013,9 @@ async function main() {
|
|
|
977
1013
|
return;
|
|
978
1014
|
}
|
|
979
1015
|
emit(flags, res, [
|
|
1016
|
+
...(res.startedElsewhere
|
|
1017
|
+
? [`NOTE: recorded starting on ${res.startedElsewhere.recorded}, replayed from ${res.startedElsewhere.here}`]
|
|
1018
|
+
: []),
|
|
980
1019
|
...(res.results ?? []).map(stepLine),
|
|
981
1020
|
actions.flowSummary(res, { withTime: false }),
|
|
982
1021
|
]);
|
|
@@ -1236,8 +1275,14 @@ async function main() {
|
|
|
1236
1275
|
'flow runs agent p50 human p50 HPI_time step_ratio turns esc',
|
|
1237
1276
|
...report.flows.map((f) => metrics.flowRow(f, { wide: true })),
|
|
1238
1277
|
'',
|
|
1239
|
-
`HPI_accuracy ${report.overall.hpi_accuracy} (${report.overall.runs}
|
|
1278
|
+
`HPI_accuracy ${report.overall.hpi_accuracy} (${report.overall.runs} measurable run(s), ` +
|
|
1240
1279
|
`${report.overall.runs - runs.filter((r) => r.completed && !r.wrong_action_taken).length} not clean)`,
|
|
1280
|
+
// A denominator that quietly shrinks is worse than one that is wrong.
|
|
1281
|
+
report.overall.runs_lost_to_device
|
|
1282
|
+
? ` ${report.overall.runs_lost_to_device} further run(s) left the denominator because the DEVICE failed,`
|
|
1283
|
+
+ ' not the code — those are unmeasured, not inaccurate (item 173):'
|
|
1284
|
+
+ `\n${(report.overall.device_causes ?? []).map((c) => ` ${c}`).join('\n')}`
|
|
1285
|
+
: null,
|
|
1241
1286
|
report.overall.hpi_time == null
|
|
1242
1287
|
? `HPI_time and HPI need a human baseline — none of ${report.overall.flows_measured} measured flow(s) has one yet.`
|
|
1243
1288
|
: `HPI_time ${report.overall.hpi_time} (harmonic mean over ${report.overall.flows_with_human_baseline} flow(s)), HPI ${report.overall.hpi}`,
|
|
@@ -1248,7 +1293,10 @@ async function main() {
|
|
|
1248
1293
|
`steps per model call ${report.overall.steps_per_call ?? '—'}`
|
|
1249
1294
|
+ (report.overall.steps_per_call
|
|
1250
1295
|
? ` — about ${(((20000 + 1700 * report.overall.steps_per_call) / report.overall.steps_per_call) / 1000).toFixed(1)}s`
|
|
1251
|
-
+ ' per step end to end
|
|
1296
|
+
+ ' per step end to end. Raise this, not the engine — but note the ~1.7s'
|
|
1297
|
+
+ ' figure for simframe\'s own work is WARM taps inside a batch: a cold app'
|
|
1298
|
+
+ ' launch is a large one-off on top, and it dominates short flows'
|
|
1299
|
+
+ ' (measured: 6.2s/step over 2 steps with no model at all).'
|
|
1252
1300
|
: ''),
|
|
1253
1301
|
flags.out ? `wrote ${flags.out}` : null,
|
|
1254
1302
|
]);
|
|
@@ -1337,13 +1385,45 @@ async function main() {
|
|
|
1337
1385
|
if (read > 0) {
|
|
1338
1386
|
verdict = named + (read < n ? ` (on the ${read} of ${n} whose reason was read)` : '');
|
|
1339
1387
|
} else if (assumed > 0) {
|
|
1340
|
-
|
|
1388
|
+
// For `verification_failed` this line used to end the story, and
|
|
1389
|
+
// it is no longer the whole truth: the verdict split below names
|
|
1390
|
+
// a faculty for some of them and takes the device's own failures
|
|
1391
|
+
// out altogether.
|
|
1392
|
+
verdict = r === 'verification_failed' && (b.verdicts?.length || b.device_total)
|
|
1393
|
+
? 'reason too coarse to steer by — see the verdict split below'
|
|
1394
|
+
: 'reason assumed, not read — no faculty can be named from these';
|
|
1341
1395
|
} else {
|
|
1342
1396
|
// Neither read nor assumed: the log predates the distinction.
|
|
1343
|
-
verdict = `${named} — but these records predate the check, so treat it as untested
|
|
1397
|
+
verdict = `${named} — but these records predate the check, so treat it as untested`
|
|
1398
|
+
+ ' (legacy, not assumed)';
|
|
1344
1399
|
}
|
|
1345
1400
|
return ` ${r.padEnd(20)} ${String(n).padStart(4)} ${verdict}`;
|
|
1346
1401
|
}),
|
|
1402
|
+
// The two things the reason alone could not say.
|
|
1403
|
+
//
|
|
1404
|
+
// Both exist because the per-reason table above was the steering wheel
|
|
1405
|
+
// and it pointed at one faculty for a class holding three, and counted
|
|
1406
|
+
// the simulator's own failures towards a perception phase.
|
|
1407
|
+
b.verdicts?.length ? '' : null,
|
|
1408
|
+
b.verdicts?.length
|
|
1409
|
+
? 'verification_failed, split by the verdict that fired'
|
|
1410
|
+
+ (b.derived?.verdicts
|
|
1411
|
+
? ` (${b.derived.verdicts} of ${b.verdicts.reduce((a, v) => a + v.count, 0)} derived from the detail text, not recorded at the time):`
|
|
1412
|
+
: ':')
|
|
1413
|
+
: null,
|
|
1414
|
+
...(b.verdicts ?? []).slice(0, 8).map((v) => {
|
|
1415
|
+
const faculty = metrics.VERDICT_FACULTY[v.name];
|
|
1416
|
+
return ` ${v.name.padEnd(20)} ${String(v.count).padStart(4)} `
|
|
1417
|
+
+ (faculty
|
|
1418
|
+
? `points at: ${faculty}`
|
|
1419
|
+
: 'no faculty follows from this verdict alone — see metrics.VERDICT_FACULTY');
|
|
1420
|
+
}),
|
|
1421
|
+
b.device_total ? '' : null,
|
|
1422
|
+
b.device_total
|
|
1423
|
+
? `${b.device_total} escalation(s) were the DEVICE, not the code — item 173, and not evidence for any faculty`
|
|
1424
|
+
+ `${b.derived?.device ? ` (${b.derived.device} derived from the detail text)` : ''}:`
|
|
1425
|
+
: null,
|
|
1426
|
+
...(b.device ?? []).slice(0, 6).map((d) => ` ${String(d.count).padStart(4)} ${d.name}`),
|
|
1347
1427
|
b.total ? '' : null,
|
|
1348
1428
|
b.total ? `avoidable ${b.avoidable}/${b.total} (${b.avoidable_escalation_rate})` : null,
|
|
1349
1429
|
// Said out loud rather than left for someone to discover: the rate is
|