simframe 0.18.0 → 0.19.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +132 -1101
- package/native/simframed/Sources/PrivateAPI/CoreSimulatorPlatform.swift +12 -6
- package/native/simframed/Sources/simframed/main.swift +18 -2
- package/package.json +1 -1
- package/scripts/article-md.mjs +111 -45
- package/scripts/bench-hpi.mjs +45 -1
- package/scripts/ci-memory.mjs +33 -6
- package/scripts/demo-gif/README.md +36 -0
- package/scripts/demo-gif/compose.swift +106 -0
- package/scripts/demo-gif/events.example.json +74 -0
- package/scripts/demo-gif/flow.json +6 -0
- package/scripts/smithery/icon.png +0 -0
- package/scripts/smithery-bundle.mjs +80 -0
- package/scripts/smithery-metadata.mjs +47 -0
- package/scripts/sync-server-version.mjs +19 -6
- package/skills/simframe/SKILL.md +7 -0
- package/src/actions.js +339 -33
- package/src/cli.js +46 -27
- package/src/device-state.js +37 -0
- package/src/index.js +144 -12
- package/src/mcp.js +2 -2
- package/src/platform/cdp.js +242 -0
- package/src/platform/index.js +28 -2
- package/src/refs.js +18 -4
- package/src/store.js +39 -0
- package/src/view.js +6 -2
- package/src/wedge.js +154 -2
package/README.md
CHANGED
|
@@ -3,152 +3,38 @@
|
|
|
3
3
|
[](https://github.com/lvlrSajjad/simframe/actions/workflows/ci.yml)
|
|
4
4
|
[](https://www.npmjs.com/package/simframe)
|
|
5
5
|
[](./LICENSE)
|
|
6
|
+
[](https://glama.ai/mcp/servers/lvlrSajjad/simframe)
|
|
6
7
|
|
|
7
|
-
**Eyes, hands and memory for
|
|
8
|
+
**Eyes, hands and memory for a coding agent driving the iOS Simulator or an
|
|
9
|
+
Android emulator.** The agent reads the screen as text with tap points, runs a
|
|
10
|
+
whole flow in one call, and remembers screens it has seen — so it stops paying
|
|
11
|
+
a screenshot, a model round trip and a re-read for every single step.
|
|
8
12
|
|
|
9
|
-
[
|
|
13
|
+

|
|
10
14
|
|
|
11
|
-
|
|
12
|
-
one is obvious:
|
|
15
|
+
## Install in ten seconds
|
|
13
16
|
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
captures mid-animation and has to look again.
|
|
17
|
-
2. **Every step is a round trip.** Tap, screenshot, reason, tap, screenshot. A
|
|
18
|
-
twelve-step flow costs twelve model turns, and the model turns cost far more
|
|
19
|
-
than the milliseconds.
|
|
20
|
-
3. **Nothing is remembered.** The same screen gets re-read and re-reasoned about
|
|
21
|
-
every single time it appears.
|
|
17
|
+
You need a Mac with Xcode (you have one if you have a simulator) and Node 18+.
|
|
18
|
+
Nothing else — no idb, no Appium, no Python.
|
|
22
19
|
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
as a native image block and 15,000–25,000 when it is not, and it does not tell
|
|
26
|
-
you what is tappable or where — you have to measure that by eye.
|
|
27
|
-
|
|
28
|
-
simframe attacks all four: a background loop keeps the newest frame warm, whole
|
|
29
|
-
flows run in one call, screens the agent has seen before are answered from
|
|
30
|
-
memory, and every answer is text with tap points in it. Nothing returns an
|
|
31
|
-
image unless you ask for one.
|
|
32
|
-
|
|
33
|
-
### It is also a study case, and the numbers are the point
|
|
34
|
-
|
|
35
|
-
Almost everything here was decided by a measurement rather than by an argument,
|
|
36
|
-
and several of those measurements **reversed a decision that had already
|
|
37
|
-
shipped**. They are written down in full, with what we expected beforehand, what
|
|
38
|
-
it cost to find out, and the mistakes made getting there — because a benchmark
|
|
39
|
-
that only records the winner teaches nothing.
|
|
40
|
-
|
|
41
|
-
| | |
|
|
42
|
-
| --- | --- |
|
|
43
|
-
| [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | the record: N, median, p95, machine, every round |
|
|
44
|
-
| [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
|
|
45
|
-
| [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
|
|
46
|
-
| [`docs/ARTICLE.md`](docs/ARTICLE.md) | the argument the whole thing adds up to |
|
|
47
|
-
|
|
48
|
-
The headline results, all free to reuse:
|
|
49
|
-
|
|
50
|
-
- **A warm transition graph removes ~27% of an agent's model round trips** —
|
|
51
|
-
isolated with a three-pass field experiment, and corroborated to three points
|
|
52
|
-
by an independent estimate. Operator knowledge is worth twice that, and no
|
|
53
|
-
tool here carries it yet. ([§14](docs/EXPERIMENTS.md))
|
|
54
|
-
- **A bigger local judge is not a better one.** `qwen3:14b` scored *lower* than
|
|
55
|
-
`qwen3:8b` — 82% against 91% — while being 79% larger and 62% slower. A
|
|
56
|
-
one-line threshold on a number already computed beat all of them at **95% and
|
|
57
|
-
zero latency**.
|
|
58
|
-
- **Every interactive element came from the accessibility tree — 72 of 72 — and
|
|
59
|
-
83% of them have no text at all.** Screen *recognition*, though, did not need
|
|
60
|
-
the tree once. The ladder splits there: semantics need a tree, identity needs
|
|
61
|
-
pixels.
|
|
62
|
-
- **Capture is ~31× faster than a screenshot** for the same work, and the
|
|
63
|
-
capture primitive itself about a thousand times faster.
|
|
64
|
-
- **Three independent reporters converged on one thing**: a refusal is cheap and
|
|
65
|
-
a confident wrong answer is expensive. Every serious bug they found was a
|
|
66
|
-
component reporting more certainty than it had.
|
|
67
|
-
|
|
68
|
-
## What changed, measured
|
|
69
|
-
|
|
70
|
-
Same four-tab navigation flow, on a real production app:
|
|
71
|
-
|
|
72
|
-
| | Before | With simframe |
|
|
73
|
-
| --- | --- | --- |
|
|
74
|
-
| Look at the screen | ~130–400 ms, blocking | **~20 ms**, already captured |
|
|
75
|
-
|
|
76
|
-
| "Did anything change?" | a full image | **~2 ms**, text only |
|
|
77
|
-
| Finding a control | read tree (~570 ms) + reason | **~1 ms** from memory |
|
|
78
|
-
| A 4-step flow, verified | 4+ model round trips | **1 call**, 3.6 s |
|
|
79
|
-
| Same flow, 3rd run | no improvement — every run is the first | **3.7 s, 4/4 from memory, 4/4 verified** |
|
|
80
|
-
| A 10-step flow | 10 turns, 10 images (~16,000 tokens at best) | **1 turn, 0 images, ~1,650 characters** |
|
|
81
|
-
| Reading a screen | an image: ~1,600 tokens, no tap points | **~330 tokens** of text, with tap points |
|
|
82
|
-
|
|
83
|
-
Every figure above is the cost inside a live process — the MCP server, or the
|
|
84
|
-
daemon answering a socket — which is how an agent actually uses simframe. A
|
|
85
|
-
one-shot `simframe` command from a shell pays about 200 ms of Node startup on
|
|
86
|
-
top, and a frame sitting on an idle screen can be older than 20 ms because the
|
|
87
|
-
capture loop throttles when nothing moves. `~20 ms` is the read, not the
|
|
88
|
-
process.
|
|
89
|
-
|
|
90
|
-
The four-tab tour, three times back to back from a cleared memory:
|
|
91
|
-
|
|
92
|
-
| Pass | Wall clock | Steps verified | Controls from memory |
|
|
93
|
-
| --- | --- | --- | --- |
|
|
94
|
-
| 1 | 10.2 s | 0/4 — nothing is known yet | **4/4** |
|
|
95
|
-
| 2 | **3.6 s** | **4/4** | **4/4** |
|
|
96
|
-
| 3 | **3.7 s** | **4/4** | **4/4** |
|
|
97
|
-
|
|
98
|
-
Every step is checked against what the same action did last time, and the run
|
|
99
|
-
records its own preconditions — which input path, which daemon, whether the
|
|
100
|
-
daemon was replaced mid-run — so a regression shows up in the measurement rather
|
|
101
|
-
than hiding inside it. Earlier versions of this table quoted 7.4 s → 3.3 s with
|
|
102
|
-
verification switched off; those numbers were measured while input was silently
|
|
103
|
-
falling back to a slower path and the capture daemon was being replaced by every
|
|
104
|
-
command, so they measured two bugs rather than the tool.
|
|
105
|
-
|
|
106
|
-
## Install
|
|
20
|
+
**Claude Code** — as a plugin, which brings the MCP server and the
|
|
21
|
+
[skill](skills/simframe/SKILL.md) that teaches the protocol in one install:
|
|
107
22
|
|
|
108
23
|
```bash
|
|
109
|
-
|
|
110
|
-
simframe
|
|
24
|
+
claude plugin marketplace add lvlrSajjad/simframe
|
|
25
|
+
claude plugin install simframe@simframe
|
|
111
26
|
```
|
|
112
27
|
|
|
113
|
-
|
|
114
|
-
from source — including `doctor`, which is why a cold `doctor` takes around 15
|
|
115
|
-
seconds and every later one takes two. It needs the Xcode command line tools,
|
|
116
|
-
which you already have if you have a simulator. Without them simframe falls
|
|
117
|
-
back to the original `simctl` loop and says so.
|
|
118
|
-
|
|
119
|
-
If more than one simulator is booted, name the one you mean — `--device=<udid>`,
|
|
120
|
-
or `export SIMFRAME_DEVICE=<udid>` once per shell. simframe refuses to choose
|
|
121
|
-
for you, because the first booted device is nobody's idea of "yours" and the
|
|
122
|
-
command that would act on it is a tap.
|
|
123
|
-
|
|
124
|
-
`doctor` checks each capability separately and tells you what you have:
|
|
125
|
-
|
|
126
|
-
```
|
|
127
|
-
ok xcrun xcrun version 72.
|
|
128
|
-
ok sips available
|
|
129
|
-
ok input driver simframed: Indigo HID
|
|
130
|
-
ok accessibility tree simframed: AXPTranslator, host-side
|
|
131
|
-
ok on-device OCR available
|
|
132
|
-
ok booted simulator iPhone 17 Pro (iOS 26.5)
|
|
133
|
-
ok capture frame #888 322x700 in 2ms (age 538ms)
|
|
134
|
-
ok sensor mode full — accessibility and OCR fused on every read (~164ms)
|
|
135
|
-
ok local supervisor none — not requested (SIMFRAME_SUPERVISOR is unset)
|
|
136
|
-
ok local planner none — not requested (SIMFRAME_PLANNER is unset)
|
|
137
|
-
```
|
|
138
|
-
|
|
139
|
-
The last three are experiments and `none` is their normal answer. See
|
|
140
|
-
[Local tiers, off by default](#local-tiers-off-by-default).
|
|
141
|
-
|
|
142
|
-
### Claude Code
|
|
28
|
+
or the MCP server alone:
|
|
143
29
|
|
|
144
30
|
```bash
|
|
145
31
|
claude mcp add --scope user simframe -- npx -y simframe mcp
|
|
146
32
|
```
|
|
147
33
|
|
|
148
|
-
|
|
149
|
-
|
|
34
|
+
Pick one, not both: two installs mean two servers driving the same device. If
|
|
35
|
+
you added it with `claude mcp add` before, `claude mcp remove simframe` first.
|
|
150
36
|
|
|
151
|
-
|
|
37
|
+
**Claude Desktop** — Settings → Developer → Edit Config, then add:
|
|
152
38
|
|
|
153
39
|
```json
|
|
154
40
|
{
|
|
@@ -158,271 +44,52 @@ registered only for the directory you ran the command in.
|
|
|
158
44
|
}
|
|
159
45
|
```
|
|
160
46
|
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
The measured cost of driving an app is not perception — warm, an
|
|
164
|
-
accessibility-only read is 85 ms and a fused read 142 ms. It is **round trips**:
|
|
165
|
-
in one instrumented run, 75% of the wall time was the agent thinking and the
|
|
166
|
-
call boundary, not simframe working. So the tools that matter most are the ones
|
|
167
|
-
that let a batch survive a problem instead of handing it back.
|
|
168
|
-
|
|
169
|
-
**Fallback selectors.** `{"tap": "Save", "or": ["Done", "Confirm"]}` — tried
|
|
170
|
-
locally in order, only an exhausted list reaching the model. Eligible after a
|
|
171
|
-
selector that did not *resolve* and nothing else, because retrying from a screen
|
|
172
|
-
you did not expect to be on is a second guess. A destructive-looking label is
|
|
173
|
-
refused as a substitute even if you list it.
|
|
174
|
-
|
|
175
|
-
**Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
|
|
176
|
-
"optional": true}` is skipped when nothing matches and runs normally when
|
|
177
|
-
something does. It exists because the opposite cost real time: a batch that
|
|
178
|
-
included a dismissal for a first-launch sheet lost **six correct steps** on the
|
|
179
|
-
next run because the sheet *did not appear*. A nag screen, a permission prompt,
|
|
180
|
-
a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
|
|
181
|
-
is the expensive outcome — a call per step instead of a call per flow. Only a
|
|
182
|
-
selector that resolved to **nothing** is absorbed: a target that is on screen
|
|
183
|
-
twice is ambiguous, not absent, and must still verify. Skipped steps are
|
|
184
|
-
reported, because "the sheet was gone" and "the sheet was dismissed" are
|
|
185
|
-
different facts.
|
|
186
|
-
|
|
187
|
-
**`{"seek": "change username", "budget": 6}`** opens containers, checks, and
|
|
188
|
-
comes back, depth first, inside a hard budget. It **acts** — opening a door
|
|
189
|
-
changes state — and it refuses to open anything that commits, abandons or
|
|
190
|
-
answers. It does not tap the target; it leaves you on the screen where the target
|
|
191
|
-
resolves.
|
|
192
|
-
|
|
193
|
-
**`{"sweep": "all", "fill": {…}}`** reads a long screen a viewport at a time and
|
|
194
|
-
fills each field while it is on screen. A form taller than the screen is only
|
|
195
|
-
knowable in pieces — the tree publishes what is rendered — and one scroll gesture
|
|
196
|
-
travels a non-deterministic distance, so finding a field and scrolling back to it
|
|
197
|
-
does not work. Sweeping does: on a real web form it filled every field in one
|
|
198
|
-
call. It detects both ends by measuring how far the *content* moved, ignoring
|
|
199
|
-
fixed chrome, which is the only reliable signal available since nothing reports a
|
|
200
|
-
scroll offset.
|
|
201
|
-
|
|
202
|
-
**`worked here before:`** puts the graph's own vocabulary in the map, most-used
|
|
203
|
-
first, rather than reporting a count. When the remembered controls are *not* on
|
|
204
|
-
the screen it says so instead, because that means two screens share one
|
|
205
|
-
fingerprint — and confident advice on a misidentified screen is how a remembered
|
|
206
|
-
label ends up pointing at a submit button.
|
|
207
|
-
|
|
208
|
-
## Local tiers, off by default
|
|
209
|
-
|
|
210
|
-
Two on-device model experiments, both `none` unless asked for, both degrading to
|
|
211
|
-
the existing matcher-then-model ladder, and CI runs with both off. They ship no
|
|
212
|
-
weights: Apple's Foundation Models framework has nothing to download, which is
|
|
213
|
-
the whole reason it clears this project's non-goal on shipping model weights.
|
|
214
|
-
|
|
215
|
-
| flag | what it does |
|
|
216
|
-
|---|---|
|
|
217
|
-
| `--sensor=ax-first` | read the accessibility tree alone (~50 ms) and pay for OCR only when a resolve fails |
|
|
218
|
-
| `--planner=apple` | order the containers `seek` opens; it cannot choose an action |
|
|
219
|
-
| `--supervisor=apple` | when a step fails, answer `wait`, `retry` or `stop` — nothing else — before the failure reaches the model |
|
|
220
|
-
|
|
221
|
-
All three are also per-call arguments on every MCP tool, because an MCP server's
|
|
222
|
-
environment is fixed when it spawns and comparing two modes inside one session
|
|
223
|
-
was otherwise impossible.
|
|
224
|
-
|
|
225
|
-
**What is measured and what is not.** The ranker: 5 of 6 top-1 on hand-written
|
|
226
|
-
cases, median 564 ms warm, and on a real exploration it went to the right region
|
|
227
|
-
in two steps where reading order wandered into version strings. `ax-first` made
|
|
228
|
-
no measurable difference to how an agent drove a real app, with one small
|
|
229
|
-
regression and one small win.
|
|
230
|
-
|
|
231
|
-
**The supervisor, on 22 labelled failures from a seeded React Native app** —
|
|
232
|
-
and the number to judge it by is round trips, not accuracy:
|
|
233
|
-
|
|
234
|
-
| | |
|
|
235
|
-
|---|---|
|
|
236
|
-
| handled locally, no round trip | **9 of 22** |
|
|
237
|
-
| escalated to Claude | 13 of 22 |
|
|
238
|
-
| a model round trip, measured in the field | 10–16 s |
|
|
239
|
-
|
|
240
|
-
So roughly 90–145 s saved on that population. Accuracy was 77 / 82 / 86% across
|
|
241
|
-
three runs of the *same* questions — it is **not deterministic**, and every
|
|
242
|
-
earlier single-run figure in this project carried that spread without reporting
|
|
243
|
-
it. All of its errors were `wait` where `stop` was right, which is the cheap
|
|
244
|
-
direction: a wrong `wait` costs a settle and a re-run, and `stop` already means
|
|
245
|
-
"hand back to Claude with the unattempted steps" rather than "give up".
|
|
246
|
-
|
|
247
|
-
**Still a bench, not the field.** These are fixtures we designed, on a testbed we
|
|
248
|
-
built, labelled by the person who then scored them. Numbers and conditions are in
|
|
249
|
-
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md); the judgements, including a phase
|
|
250
|
-
cancelled by its own measurement, are in [`docs/DECISIONS.md`](docs/DECISIONS.md).
|
|
251
|
-
|
|
252
|
-
**And the measurements that changed our minds** are in
|
|
253
|
-
[`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md), with what we expected beforehand
|
|
254
|
-
written down beside each — including the capacity comparison the supervisor
|
|
255
|
-
design was assumed to make unnecessary. Asked the same 22 situations, three
|
|
256
|
-
times each:
|
|
257
|
-
|
|
258
|
-
| arm | accuracy | median | deterministic |
|
|
259
|
-
|---|---|---|---|
|
|
260
|
-
| always the commonest answer | 55% | — | — |
|
|
261
|
-
| **`stillMs > 3000ms`, no model at all** | **95%** | **0 ms** | yes |
|
|
262
|
-
| Apple Foundation Models (~3B, on-device) | 77 / 82 / 86% | ~640 ms | **no** |
|
|
263
|
-
| `qwen3:8b` via Ollama (4-bit, 5.2 GB) | **91%** | 919 ms | yes |
|
|
264
|
-
| `qwen3:14b` via Ollama (4-bit, 9.3 GB) | 82% | 1,489 ms | yes |
|
|
265
|
-
|
|
266
|
-
Three results, and the third is the one that matters. The larger model scored
|
|
267
|
-
*lower* than the smaller one. A free threshold on a number the daemon already
|
|
268
|
-
computes beat all three, on a population it was never fitted to. And **the
|
|
269
|
-
accuracy ranking inverts the safety ranking**: every arm errs in one direction
|
|
270
|
-
only, Apple's errors are all `wait` where `stop` was right and both Qwen arms'
|
|
271
|
-
are all `stop` where `wait` was right — and a wrong `stop` abandons a plan that
|
|
272
|
-
would have worked, while a wrong `wait` costs a settle. A comparison reporting
|
|
273
|
-
only the percentages would have recommended the wrong model.
|
|
274
|
-
|
|
275
|
-
The Ollama arm is an **experiment, not a recommendation**: off unless named,
|
|
276
|
-
no weights shipped, no dependency added, and every arm reads the same briefing
|
|
277
|
-
out of `native/supervise.swift` so no arm is answering a different question.
|
|
278
|
-
|
|
279
|
-
**Two further results, both negative, both worth more than the table.** A
|
|
280
|
-
*cascade* — free rule first, model where it is unsure, Claude after that — was
|
|
281
|
-
worse at every abstention band we tried (95% for the rule alone; 91%, 86%, 82%
|
|
282
|
-
as more was handed to the model). And giving the model a fourth word, `abstain`,
|
|
283
|
-
cost the Apple arm about a third of its accuracy (77/82/86% → 45/50/55%) **while
|
|
284
|
-
it never used the word once**. One added paragraph, nothing else changed. So the
|
|
285
|
-
fourth word ships off, and "add an abstain token" became "find a judge that will
|
|
286
|
-
use one".
|
|
287
|
-
|
|
288
|
-
The supervisor's whole vocabulary is three words on purpose. It cannot invent a
|
|
289
|
-
step, skip one, substitute a target or continue past an unexpected screen — not
|
|
290
|
-
because a threshold forbids it but because those are not answers it can give.
|
|
291
|
-
That constraint replaced an earlier version of the same idea that was given
|
|
292
|
-
latitude over *what* to open and pressed a button labelled "YES, THIS FIXED MY
|
|
293
|
-
PROBLEM" in a live app.
|
|
294
|
-
|
|
295
|
-
## Capabilities are independent
|
|
296
|
-
|
|
297
|
-
Each layer works without the ones above it, and `doctor` tells you which you
|
|
298
|
-
have. **Observation needs nothing but Xcode.**
|
|
299
|
-
|
|
300
|
-
| Capability | Needs | Without it |
|
|
301
|
-
| --- | --- | --- |
|
|
302
|
-
| Watch the screen, wait, recall | nothing extra | — |
|
|
303
|
-
| Read labels + coordinates from pixels | `swiftc` (Xcode CLT) | falls back to the accessibility tree alone |
|
|
304
|
-
| Tap, type, swipe | nothing extra | simframe observes but cannot touch |
|
|
305
|
-
| Accessibility tree | nothing extra | OCR alone still yields labels and coordinates |
|
|
306
|
-
|
|
307
|
-
`simframe doctor` names which engine is carrying each capability, per device.
|
|
308
|
-
**Nothing beyond Xcode is required.** Capture, input, text recognition and the
|
|
309
|
-
accessibility tree all run in-process, in one daemon.
|
|
310
|
-
|
|
311
|
-
[`idb`](https://fbidb.io) is still accepted as a fallback for input and for the
|
|
312
|
-
tree, for a machine where the daemon cannot run — and `SIMFRAME_AX_DRIVER=idb`
|
|
313
|
-
forces the tree back onto it, which is the escape hatch if an Xcode upgrade
|
|
314
|
-
breaks the host-side path.
|
|
315
|
-
|
|
316
|
-
```bash
|
|
317
|
-
# optional fallback, not a requirement
|
|
318
|
-
brew tap facebook/fb && brew install idb-companion && pipx install fb-idb
|
|
319
|
-
```
|
|
320
|
-
|
|
321
|
-
## Android
|
|
322
|
-
|
|
323
|
-
simframe's second backend drives an Android emulator with the same commands, the
|
|
324
|
-
same screen map and the same memory as a simulator. Everything above the
|
|
325
|
-
platform boundary — the frame store, settle, the structural fingerprint, the
|
|
326
|
-
screen map, refs and the transition graph — runs on it unmodified, because the
|
|
327
|
-
boundary hands it frames and nothing above it knows what a simulator is.
|
|
47
|
+
**Cursor** — [](https://cursor.com/en/install-mcp?name=simframe&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsInNpbWZyYW1lIiwibWNwIl19)
|
|
48
|
+
or add the same `simframe` block to `~/.cursor/mcp.json`.
|
|
328
49
|
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
| Watch the screen, wait, recall | yes | frames at **41 ms** through the emulator console, host-side — no adb in the capture path |
|
|
332
|
-
| Read labels + coordinates from pixels | yes | the same Vision OCR + CV, off the same PNG |
|
|
333
|
-
| Screen map, refs, screen memory, the graph | yes | unchanged above the boundary |
|
|
334
|
-
| Tap, type, swipe, keys | yes | the console's `event mouse` as a real down/move/up, `event text` for characters, `input keyevent` for keys |
|
|
335
|
-
| Clipboard, and `paste` into a field | yes | the emulator's gRPC `setClipboard`, over `node:http2`, no dependency, then `KEYCODE_PASTE` to deliver it |
|
|
336
|
-
| List/resolve devices, launch, terminate, open a URL, permissions | yes | `adb`, with the permission state read back off the device |
|
|
337
|
-
| Accessibility tree | **not available (OCR + CV only)** | `uiautomator dump` costs **2,012 ms** a read, against 45 ms for the iOS tree. See [`docs/DEFERRED.md`](docs/DEFERRED.md) |
|
|
338
|
-
| A launch confirmed to have reached the front | **not available (the launch is not checked)** | iOS compares the pid `simctl launch` printed against the pid the device reports as frontmost, in **2–5 ms**. Nothing here reports either; `am start` fronts synchronously, which is why it has not bitten — but that is not a check. See [`docs/DEFERRED.md`](docs/DEFERRED.md) |
|
|
50
|
+
**Codex, Windsurf, Zed, VS Code, anything else that speaks MCP** — the same
|
|
51
|
+
JSON block, wherever that client keeps its `mcpServers`.
|
|
339
52
|
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
simframe do --device=emulator-5554 flow.json
|
|
345
|
-
```
|
|
53
|
+
Then boot a simulator and ask the agent for something: *"open my app and tap
|
|
54
|
+
through the signup flow"*, *"is the list on the Orders tab loading?"*, *"why
|
|
55
|
+
does the Save button do nothing?"*. The first call builds a small Swift daemon
|
|
56
|
+
from source — about 15 seconds, once.
|
|
346
57
|
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
tap a simulator while you were driving an emulator, and acting on the wrong
|
|
350
|
-
device is worse than refusing. So a command with no device names both and stops.
|
|
351
|
-
`--device` answers it per command; `SIMFRAME_DEVICE` answers it per shell:
|
|
58
|
+
The CLI is the same thing without an agent in the loop, and the cheapest way
|
|
59
|
+
to check the install:
|
|
352
60
|
|
|
353
61
|
```bash
|
|
354
|
-
|
|
355
|
-
simframe
|
|
62
|
+
npm install -g simframe
|
|
63
|
+
simframe doctor # every layer, and which engine carries it
|
|
64
|
+
simframe ui # the current screen as a numbered text map
|
|
356
65
|
```
|
|
357
66
|
|
|
358
|
-
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
installed onto your device. The perception ladder was built so a missing tier
|
|
362
|
-
degrades rather than fails, and this is exactly that case: OCR and CV yield
|
|
363
|
-
labels and coordinates on Android today, and a tap by label works without a tree
|
|
364
|
-
at all. `simframe doctor` reports the tier as `optional` with that number, so
|
|
365
|
-
the gap is visible rather than silent, and the criteria for revisiting it are in
|
|
366
|
-
`docs/DEFERRED.md` under **Phase 8b**.
|
|
367
|
-
|
|
368
|
-
**What the missing tree costs, measured rather than hand-waved.** Screen
|
|
369
|
-
*identity* is weaker on Android than on iOS, and specifically so. Tokens per
|
|
370
|
-
screen, and where they come from:
|
|
371
|
-
|
|
372
|
-
| Screen | Tokens | Regions | Roles | Chrome labels |
|
|
373
|
-
| --- | --- | --- | --- | --- |
|
|
374
|
-
| launcher | **1** | nav-bar 1 | text 1 | 0 |
|
|
375
|
-
| Settings root | 9 | content 9 | text 9 | 0 |
|
|
376
|
-
| example.com in Chrome | 6 | content 3, nav-bar 3 | text 6 | 0 |
|
|
377
|
-
|
|
378
|
-
Every token has role `text`, because without a tree nothing infers a button from
|
|
379
|
-
a rectangle reliably enough to say so, and no screen here carries a chrome label
|
|
380
|
-
at all. So on Android a screen is recognised by the geometry of its text, which
|
|
381
|
-
is thinner and noisier than the iOS mix of roles, chrome labels and geometry.
|
|
382
|
-
Flows still work; screen *memory* is doing more guessing, and that is the honest
|
|
383
|
-
cost of the tier being absent.
|
|
67
|
+
If more than one simulator is booted, name the one you mean with
|
|
68
|
+
`--device=<udid>` or `export SIMFRAME_DEVICE=<udid>`; simframe refuses to guess,
|
|
69
|
+
because the command that would act on the wrong one is a tap.
|
|
384
70
|
|
|
385
|
-
|
|
386
|
-
out of identity, which would be a strict improvement there — is not available:
|
|
387
|
-
it would leave Settings' root with zero tokens, and zero tokens is no identity
|
|
388
|
-
at all. See [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
71
|
+
## What it changes
|
|
389
72
|
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
address.
|
|
397
|
-
|
|
398
|
-
The emulator's own gRPC surface was checked for anything tree-shaped and has
|
|
399
|
-
nothing: 43 RPCs for sensors, input, screenshots and VM state, and no notion of
|
|
400
|
-
a view. That question is settled, not open. The same surface is what carries the
|
|
401
|
-
clipboard.
|
|
402
|
-
|
|
403
|
-
## The tools
|
|
73
|
+
An agent driving a simulator is slow for three reasons, and only the first is
|
|
74
|
+
obvious. Every look is a wait (`simctl io screenshot` costs ~130 ms, paid on
|
|
75
|
+
every glance). Every step is a model round trip. And nothing is remembered, so
|
|
76
|
+
the same screen is re-read and re-reasoned about each time it appears. There
|
|
77
|
+
is a fourth that is pure waste: an image is the most expensive way to ask what
|
|
78
|
+
is on screen, and it still does not say what is tappable.
|
|
404
79
|
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
80
|
+
| | Without simframe | With simframe |
|
|
81
|
+
| --- | --- | --- |
|
|
82
|
+
| Look at the screen | ~130–400 ms screenshot, blocking | **~20 ms**, already captured |
|
|
83
|
+
| "Did anything change?" | a full image | **~2 ms**, text only |
|
|
84
|
+
| Reading a screen | an image, ~1,600 tokens, no tap points | **~330 tokens** of text, with tap points |
|
|
85
|
+
| A 10-step flow | 10 turns, 10 images | **1 turn, 0 images** |
|
|
86
|
+
| Same flow, again | every run is the first | replayed from memory, every step verified, **zero model calls** |
|
|
408
87
|
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
| `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
|
|
413
|
-
| `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
|
|
414
|
-
| `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
|
|
415
|
-
| `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
|
|
416
|
-
| `sim_find` | Resolve an intent to one control, without acting on it. |
|
|
417
|
-
| `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
|
|
418
|
-
| `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
|
|
419
|
-
| `sim_wait` | Waits for the screen to change *and then* settle. |
|
|
420
|
-
| `sim_look` | **The only tool that returns an image**, capped at 1024 px. For layout, colour, spacing — questions text cannot answer. |
|
|
421
|
-
| `sim_recall` · `sim_strip` | Look backwards: a text timeline of what happened, or recent frames tiled into one image. |
|
|
422
|
-
| `sim_storage` | **What the app believes**, as opposed to what it drew: its `UserDefaults` and, for React Native, its `AsyncStorage`. Reads the data container off disk, so it answers on a device that is **not running**. |
|
|
423
|
-
| `sim_capture` · `sim_devices` | Manage capture loops; list simulators. |
|
|
88
|
+
Measured on an iPhone 17 Pro simulator, iOS 26.5, Apple Silicon, against a real
|
|
89
|
+
production app; the full tables with N, median and p95 are in
|
|
90
|
+
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
424
91
|
|
|
425
|
-
|
|
92
|
+
This is what the agent gets back instead of a picture:
|
|
426
93
|
|
|
427
94
|
```
|
|
428
95
|
iPhone 17 Pro · 402x874pt · screen a1b2c3d4 "Inbox" (known, 3 known exits)
|
|
@@ -437,745 +104,109 @@ content:
|
|
|
437
104
|
tab-bar:
|
|
438
105
|
#6 text 62,835 Inbox
|
|
439
106
|
#7 text 201,835 Settings
|
|
107
|
+
next: settled; nothing ambiguous — chain the next steps in one sim_do without looking again.
|
|
440
108
|
```
|
|
441
109
|
|
|
442
|
-
Region
|
|
443
|
-
|
|
444
|
-
|
|
445
|
-
|
|
446
|
-
different screen it refuses rather than tapping whatever now sits there.
|
|
447
|
-
|
|
448
|
-
`= something` is what the control *contains*, from the accessibility tree, and
|
|
449
|
-
`~ something` is what OCR read off the pixels. Both are printed, and where they
|
|
450
|
-
disagree that is the point: one is authoritative and the other is what is
|
|
451
|
-
actually on screen, and a field mid-edit can legitimately differ. A row with no
|
|
452
|
-
`=` is a control that reports no value, not an empty one.
|
|
453
|
-
|
|
454
|
-
When the elements were recalled from screen memory rather than looked at just
|
|
455
|
-
now, the header says so and how long ago — `elements recalled from 41s ago —
|
|
456
|
-
pass refresh for what is there now`. Identity is cached on purpose, because a
|
|
457
|
-
list with new rows is the same screen; contents are exactly what changes without
|
|
458
|
-
the screen changing, so the age is worth seeing.
|
|
459
|
-
|
|
460
|
-
Three ways to name a control, anywhere one is named:
|
|
461
|
-
|
|
462
|
-
| | |
|
|
463
|
-
| --- | --- |
|
|
464
|
-
| `"Save"` · `the Assets tab` · `back` | **start here** — resolved by intent: verbs, typos, synonyms, icon-only controls by their common name |
|
|
465
|
-
| `#3` | the number the map gave it. Cheap and exact, but only within the round trip that numbered it |
|
|
466
|
-
| `@120,400` | raw point coordinates. Last resort: it cannot tell you it missed. |
|
|
467
|
-
|
|
468
|
-
The order is deliberate and it used to be the other way round. Four peer rounds
|
|
469
|
-
reported that intent resolution worked every time while refs renumbered
|
|
470
|
-
underneath them, so a table that led with `#3` and called it "unambiguous" was
|
|
471
|
-
recommending the more brittle of the two.
|
|
472
|
-
|
|
473
|
-
## Baselines: the thing to understand
|
|
474
|
-
|
|
475
|
-
Every change question is really "changed **since when**?" — and the answer is
|
|
476
|
-
almost never "since the previous frame". A UI transition is over in about 700 ms,
|
|
477
|
-
so comparing consecutive frames tells a caller that polls every few seconds
|
|
478
|
-
"nothing changed", even though the screen is completely different from when it
|
|
479
|
-
last looked.
|
|
480
|
-
|
|
481
|
-
So simframe compares against **the last frame you observed**. Over MCP that is
|
|
482
|
-
automatic. From the CLI, capture a baseline before you act:
|
|
483
|
-
|
|
484
|
-
```bash
|
|
485
|
-
H=$(simframe mark)
|
|
486
|
-
# ...tap, launch, navigate...
|
|
487
|
-
simframe wait --since=$H # change, then settle
|
|
488
|
-
simframe state --since=$H # what moved, as text
|
|
489
|
-
```
|
|
490
|
-
|
|
491
|
-
The same applies to waiting. `--mode=settle` (the default) waits for a change and
|
|
492
|
-
*then* for stillness, because a bare "wait until stable" called in the moment
|
|
493
|
-
before an animation starts will correctly, and uselessly, return immediately.
|
|
494
|
-
|
|
495
|
-
## Screen memory
|
|
496
|
-
|
|
497
|
-
An accessibility tree is a promise apps do not always keep. In testing against a
|
|
498
|
-
real production app, its custom tab bar published **no children at all**, its
|
|
499
|
-
icon buttons carried unreadable private-use glyphs, and its React Native text
|
|
500
|
-
inputs were **absent from the tree entirely** — the controls used most were
|
|
501
|
-
exactly the ones that could not be tapped by name.
|
|
502
|
-
|
|
503
|
-
So simframe reads the screen two ways and remembers the result:
|
|
504
|
-
|
|
505
|
-
- **Accessibility** gives real hit targets, types and enabled state.
|
|
506
|
-
- **On-device OCR** (Apple's Vision, ~290 ms, no model round trip) gives every
|
|
507
|
-
label a person can actually see, with coordinates.
|
|
508
|
-
- The merge is keyed by a **structural fingerprint**, so the next visit is a
|
|
509
|
-
file read. What that fingerprint is, and why it is not a pixel hash, is below.
|
|
510
|
-
|
|
511
|
-
```
|
|
512
|
-
first visit to a screen ~1000 ms read tree + OCR, store the map
|
|
513
|
-
every visit after that ~1 ms look it up
|
|
514
|
-
```
|
|
515
|
-
|
|
516
|
-
OCR is also more accurate than measuring by eye. On one tab bar the first tab
|
|
517
|
-
centre sat at x=62, not the x=40 an even five-way split predicts — a silent
|
|
518
|
-
mis-tap on every attempt.
|
|
519
|
-
|
|
520
|
-
Two details that matter:
|
|
521
|
-
|
|
522
|
-
- **Containers do not absorb their contents.** A tab bar encloses all five tab
|
|
523
|
-
labels but is not any of them, so the merge only combines an element with text
|
|
524
|
-
of comparable size.
|
|
525
|
-
- **Ambiguity is reported, not guessed.** A word that is both a screen title and
|
|
526
|
-
a tab returns an error listing both with coordinates, because silently tapping
|
|
527
|
-
the title looks exactly like nothing happening.
|
|
110
|
+
Region, type, label, contents, a tap point, and a number the next call can use
|
|
111
|
+
as a selector. `= value` is what the accessibility tree says the control
|
|
112
|
+
contains; `~ value` is what OCR read off the pixels. Where they disagree, that
|
|
113
|
+
is the point.
|
|
528
114
|
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
"Did this move?" and "is this the same screen?" look like one question and are
|
|
532
|
-
not. simframe answers them separately, and getting that wrong was the single
|
|
533
|
-
most expensive mistake in its development.
|
|
534
|
-
|
|
535
|
-
**Change and settle** are questions about pixels, so a pixel hash answers them.
|
|
536
|
-
The frame hash changes whenever any pixel group changes — a clock digit, one new
|
|
537
|
-
row — which is exactly right for "did anything happen?" and useless as a key for
|
|
538
|
-
"have I been here before?". For change detection there is a layout hash: status
|
|
539
|
-
bar cropped, difference hash over a 12×24 grid.
|
|
540
|
-
|
|
541
|
-
A mean-threshold hash was tried first and was actively dangerous: low-contrast
|
|
542
|
-
screens collapsed onto identical values, so unrelated screens matched at distance
|
|
543
|
-
0 and taps landed on the wrong control. The difference hash fixed that.
|
|
544
|
-
|
|
545
|
-
**Identity is not a question about pixels**, and this is the part that took three
|
|
546
|
-
attempts. Content *is* pixels: a list whose rows changed drifts as far as a
|
|
547
|
-
different screen does. Measured on a real app, same-screen revisits reached 62
|
|
548
|
-
bits against a different-screen floor of 74 — overlapping, with no threshold
|
|
549
|
-
available to choose. An earlier calibration had suggested a comfortable margin
|
|
550
|
-
(0–4 against 77–113), but it was measured on screens whose content happened to be
|
|
551
|
-
stable and did not survive contact with a real list.
|
|
552
|
-
|
|
553
|
-
So identity is **structural**. The fingerprint is built from element roles,
|
|
554
|
-
frames quantised to a 24 px grid, the region each element sits in, and repeated
|
|
555
|
-
siblings bucketed as "one" or "many" rather than counted. Deliberately included:
|
|
556
|
-
the labels of chrome elements only — nav title, tab labels, toolbar buttons —
|
|
557
|
-
because two list screens with identical structure are told apart by their title
|
|
558
|
-
and nothing else. Deliberately excluded: all content text and values, the status
|
|
559
|
-
bar, and the keyboard region when a keyboard is up.
|
|
560
|
-
|
|
561
|
-
It does not depend on the accessibility tree. Fingerprinting from OCR boxes
|
|
562
|
-
alone, with the tree discarded entirely, still separates screens — different
|
|
563
|
-
screens ceiling 0.35 against the same threshold.
|
|
115
|
+
## The tools
|
|
564
116
|
|
|
565
|
-
| |
|
|
117
|
+
| Tool | What it does |
|
|
566
118
|
| --- | --- |
|
|
567
|
-
|
|
|
568
|
-
|
|
|
569
|
-
|
|
570
|
-
|
|
571
|
-
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
|
|
580
|
-
|
|
581
|
-
|
|
582
|
-
|
|
583
|
-
|
|
584
|
-
|
|
585
|
-
|
|
586
|
-
|
|
587
|
-
Failing to recognise a screen you have seen is harmless — it rebuilds the map and
|
|
588
|
-
taps correctly. Matching the *wrong* screen taps the wrong control. The threshold
|
|
589
|
-
is set to err toward the first.
|
|
590
|
-
|
|
591
|
-
## Navigating by memory
|
|
592
|
-
|
|
593
|
-
Once simframe knows which screens exist and which action leads from one to the
|
|
594
|
-
next, getting somewhere is a search over known edges rather than a question for a
|
|
595
|
-
model:
|
|
119
|
+
| `sim_ui` | **Start here.** The screen as the numbered text map above. |
|
|
120
|
+
| `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. Fallback selectors, optional steps, and a bounded `seek` let a batch survive a surprise instead of handing it back. |
|
|
121
|
+
| `sim_state` | Has anything changed since your last look, and which regions moved. |
|
|
122
|
+
| `sim_goto` · `sim_flow_run` | Walk to a screen simframe has been to before; replay a saved flow with **zero model calls**. |
|
|
123
|
+
| `sim_find` · `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. |
|
|
124
|
+
| `sim_launch` · `sim_open_url` · `sim_permission` | Launch (confirmed to have reached the front), open a deep link, grant a privacy permission without tapping the alert. |
|
|
125
|
+
| `sim_look` | **The only tool that returns an image**, capped at 1024 px — for layout, colour and spacing. |
|
|
126
|
+
| `sim_recall` · `sim_strip` · `sim_storage` | What happened in the last minute, as text or as a contact sheet; what the app *saved* (`UserDefaults`, `AsyncStorage`), even on a device that is shut down. |
|
|
127
|
+
| `sim_wait` · `sim_capture` · `sim_devices` | Wait for change then settle; manage capture; list simulators and emulators. |
|
|
128
|
+
|
|
129
|
+
Three ways to name a control, anywhere one is named: `"Save"` (resolved by
|
|
130
|
+
intent — verbs, typos, synonyms, icon-only controls by their common name),
|
|
131
|
+
`#3` (the number the map gave it, valid for that screen only), or `@120,400`
|
|
132
|
+
(raw points, last resort). A label that matches the destructive vocabulary —
|
|
133
|
+
Delete, Pay, Send, Sign out — is never substituted for or guessed at.
|
|
134
|
+
|
|
135
|
+
A [Claude Code skill](skills/simframe/SKILL.md) ships in the package and
|
|
136
|
+
teaches the CLI path, which is cheaper still: a few lines of output instead of
|
|
137
|
+
a tool result, and no MCP schema in context until a tool is used.
|
|
596
138
|
|
|
597
|
-
|
|
598
|
-
simframe screens # what this device has learned
|
|
599
|
-
simframe goto invoices # walk there, verifying every step
|
|
600
|
-
```
|
|
601
|
-
|
|
602
|
-
And when the screen and the behaviour disagree, the question is usually not
|
|
603
|
-
about the screen at all:
|
|
604
|
-
|
|
605
|
-
```bash
|
|
606
|
-
simframe storage # apps with a data container
|
|
607
|
-
simframe storage com.example.myapp # what that app saved
|
|
608
|
-
```
|
|
609
|
-
|
|
610
|
-
`sim_ui` says what is drawn; `sim_storage` says what the app believes. It reads
|
|
611
|
-
the data container straight off the host filesystem, which means it works on a
|
|
612
|
-
device that is **shut down** — `simctl` cannot do this at all, on any of its own
|
|
613
|
-
paths, once a device stops running.
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
Measured on a four-tab tour, `goto` plans and walks three-step routes with every
|
|
617
|
-
step verified and no model call. It fails rather than guesses: an unknown
|
|
618
|
-
destination, a query matching two screens equally, or no path of known edges all
|
|
619
|
-
report themselves instead of tapping hopefully.
|
|
620
|
-
|
|
621
|
-
Flows work the same way and refuse to save if any step went unverified —
|
|
622
|
-
replaying a recording of something that may not have worked just reproduces the
|
|
623
|
-
doubt.
|
|
624
|
-
|
|
625
|
-
```bash
|
|
626
|
-
simframe flow save checkout ./checkout.json
|
|
627
|
-
simframe flow run checkout
|
|
628
|
-
```
|
|
629
|
-
|
|
630
|
-
### What the memory is worth, isolated
|
|
631
|
-
|
|
632
|
-
Three agent sessions drove the same task family on the same production app. The
|
|
633
|
-
third existed only to hold a variable still: pass 2 beat pass 1 by twenty-one
|
|
634
|
-
model round trips, but it had a warm graph *and* an operator who had already
|
|
635
|
-
driven the app once, so the gap could belong to either.
|
|
636
|
-
|
|
637
|
-
| pass | operator | graph | round trips |
|
|
638
|
-
| --- | --- | --- | --- |
|
|
639
|
-
| 1 | fresh | cold | **33** |
|
|
640
|
-
| 3 | fresh | **warm** | **24** |
|
|
641
|
-
| 2 | experienced | warm | **12** |
|
|
642
|
-
|
|
643
|
-
**The graph is worth ~27%** of the round trips. Operator knowledge is worth the
|
|
644
|
-
rest — half of everything left after the graph had taken its share. A second
|
|
645
|
-
reporter, unable to isolate the graph, derived **~23%** from escalation rate per
|
|
646
|
-
step without knowing that number.
|
|
647
|
-
|
|
648
|
-
One app, one operator per pass, and round trips are not seconds. What it settles
|
|
649
|
-
is the direction: the memory is real, it is the smaller half, and the larger half
|
|
650
|
-
is what an operator learns and no tool here carries between sessions yet. Full
|
|
651
|
-
working in [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) §14.
|
|
652
|
-
|
|
653
|
-
## Does this work on *your* app?
|
|
654
|
-
|
|
655
|
-
Nothing in simframe is written for a particular app. What varies between apps is
|
|
656
|
-
how much of the accessibility tree exists, and simframe is built to degrade
|
|
657
|
-
rather than fail:
|
|
658
|
-
|
|
659
|
-
- **Good tree** → tap by label, batch aggressively, everything just works.
|
|
660
|
-
- **Partial tree** (custom tab bars, icon buttons) → OCR fills the gaps; you tap
|
|
661
|
-
by the visible text instead.
|
|
662
|
-
- **No tree at all** → OCR alone still yields labels and coordinates.
|
|
139
|
+
## Android
|
|
663
140
|
|
|
664
|
-
|
|
665
|
-
|
|
666
|
-
|
|
141
|
+
The second backend drives an Android emulator with the same commands, the same
|
|
142
|
+
screen map and the same memory. Frames arrive at ~41 ms through the emulator
|
|
143
|
+
console; tap, swipe, text, keys and the clipboard work; `simframe devices`
|
|
144
|
+
lists emulators next to simulators. What Android does not have is an
|
|
145
|
+
accessibility tree, so its screens are read by OCR and CV alone — and
|
|
146
|
+
`simframe doctor` says so rather than pretending. Details, costs and the
|
|
147
|
+
reasoning are in [the guide](docs/GUIDE.md#android).
|
|
667
148
|
|
|
668
|
-
|
|
669
|
-
back chevron — is listed with its coordinates and tappable by `#ref` rather than
|
|
670
|
-
by name, and the map counts how many such controls are on the screen. A view the
|
|
671
|
-
app never declared accessible at all is invisible to any accessibility tree, ours
|
|
672
|
-
included, and the same line says so: a count of zero on a screen that has one
|
|
673
|
-
would be the more expensive answer.
|
|
149
|
+
## How it works, briefly
|
|
674
150
|
|
|
675
|
-
|
|
676
|
-
|
|
677
|
-
|
|
678
|
-
|
|
679
|
-
|
|
680
|
-
|
|
681
|
-
so: `settled after 63ms (a 13x13 region is still animating)`. A settle that
|
|
682
|
-
times out names where the movement is instead. Refusing to settle on any
|
|
683
|
-
animation would be right for a spinner and wrong for a blinking cursor, so the
|
|
684
|
-
disagreement is reported rather than resolved by a guessed threshold.
|
|
151
|
+
One Swift daemon per device reads the simulator's framebuffer straight off its
|
|
152
|
+
IOSurface when the display reports damage, runs Apple's Vision OCR on the same
|
|
153
|
+
surface, reads the accessibility tree host-side, and sends taps as real
|
|
154
|
+
down/move/up sequences over the HID channel — no PNG, no file, no spawned
|
|
155
|
+
process. Frames land in `~/.simframe/<udid>/`, input goes over a `0600` Unix
|
|
156
|
+
socket, and the MCP server and CLI are thin clients over both.
|
|
685
157
|
|
|
686
|
-
|
|
687
|
-
|
|
688
|
-
|
|
689
|
-
|
|
158
|
+
Screen memory keys a merged element map by a structural fingerprint, so a
|
|
159
|
+
screen seen before is a file read. A transition graph records which action led
|
|
160
|
+
from which screen to which, so `sim_goto` plans a route and a saved flow
|
|
161
|
+
replays with no model in the loop.
|
|
690
162
|
|
|
691
|
-
|
|
163
|
+
Every layer degrades rather than fails, and never quietly: `simframe doctor`
|
|
164
|
+
reports each one and `--strict` turns any downgrade into a non-zero exit. The
|
|
165
|
+
private frameworks the daemon links are documented, with evidence, in
|
|
166
|
+
[`docs/PRIVATE_API.md`](docs/PRIVATE_API.md).
|
|
692
167
|
|
|
693
|
-
|
|
168
|
+
## Read on
|
|
694
169
|
|
|
695
170
|
| | |
|
|
696
171
|
| --- | --- |
|
|
697
|
-
|
|
|
698
|
-
|
|
|
699
|
-
|
|
|
700
|
-
|
|
|
701
|
-
|
|
|
702
|
-
|
|
|
703
|
-
| Tap (70 ms hold / 10 ms hold) | 76 ms / 13 ms |
|
|
704
|
-
| Text recognition, in-process | **~174 ms** |
|
|
705
|
-
| Text recognition, via PNG + helper (fallback) | ~555 ms |
|
|
706
|
-
| Accessibility tree read, in-process | **~45 ms** |
|
|
707
|
-
| Accessibility tree read, via idb (fallback) | ~203 ms |
|
|
708
|
-
| Screen map: first visit / remembered | ~305 ms / **~1 ms** |
|
|
709
|
-
| CPU | 1.1 % idle · 3.1 % active |
|
|
710
|
-
| Frame memory | ~60 s of screen, ~2.7 MB |
|
|
711
|
-
|
|
712
|
-
Reproduce all of it with `npm run bench`, which prints the same table against
|
|
713
|
-
your machine. Full detail, including the measurement traps, is in
|
|
714
|
-
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
715
|
-
|
|
716
|
-
### Wall clock per step — where the time really goes
|
|
717
|
-
|
|
718
|
-
Every number above is microscopic next to the one that decides how fast this
|
|
719
|
-
feels, and it took two field reports to see it. Per-step wall clock is
|
|
720
|
-
|
|
721
|
-
```
|
|
722
|
-
(model round trip + simframe work × n) / n for n steps in one call
|
|
723
|
-
```
|
|
724
|
-
|
|
725
|
-
**Compare like with like, which this table used not to.** A human tester's
|
|
726
|
-
1.95 s per step is the *whole* loop — look at the screen, decide what to do,
|
|
727
|
-
do it. Any row that does not include a decision is not comparable to it.
|
|
728
|
-
|
|
729
|
-
| | perceives | decides | acts | per step |
|
|
730
|
-
| --- | --- | --- | --- | --- |
|
|
731
|
-
| **a human tester, measured** | yes | yes | yes | **1.95 s** |
|
|
732
|
-
| **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
|
|
733
|
-
| simframe, batch of 4 | yes | yes | yes | ~6.7 s |
|
|
734
|
-
| simframe, one model call per step | yes | yes | yes | ~21.7 s |
|
|
735
|
-
| — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
|
|
736
|
-
| — a saved flow replayed | yes | **no** | yes | 1.98 s |
|
|
737
|
-
|
|
738
|
-
**The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
|
|
739
|
-
a human**, and about eleven times when a failure forces one model call per step.
|
|
740
|
-
|
|
741
|
-
The two indented rows are the ones this README used to lead with, and both are
|
|
742
|
-
category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
|
|
743
|
-
taken out — the thinking is the model round trip, which is most of the clock.
|
|
744
|
-
And a replay *decides nothing*: it is a recording being played back, so its fair
|
|
745
|
-
counterpart is a human repeating a flow they have memorised, who would be well
|
|
746
|
-
under 1.95 s. Replay against a human working something out for the first time is
|
|
747
|
-
a rehearsal measured against a first attempt.
|
|
748
|
-
|
|
749
|
-
What the two rows do establish, and it is the finding that reordered this
|
|
750
|
-
project: **there is nothing left to win inside the engine.** 1.7 s is small
|
|
751
|
-
beside a 20 s round trip, so making perception or input faster buys single-digit
|
|
752
|
-
percentages. The only variable that matters is `n` — how many steps one decision
|
|
753
|
-
covers. Every improvement here has come from raising it, not from faster code.
|
|
754
|
-
|
|
755
|
-
### One more term: the launch is a fixed cost
|
|
756
|
-
|
|
757
|
-
Measured on the benchmark suite with **no model in the loop at all**:
|
|
758
|
-
|
|
759
|
-
| route | steps | agent | per step | human | per step |
|
|
760
|
-
| --- | --- | --- | --- | --- | --- |
|
|
761
|
-
| contacts-kate-bell | 2 | 12419 ms | **6.2 s** | 4300 ms | 2.15 s |
|
|
762
|
-
| settings-larger-text | 4 | 15281 ms | 3.8 s | 7799 ms | 1.95 s |
|
|
763
|
-
|
|
764
|
-
2.9× a human on the short route with nothing thinking, which does not fit
|
|
765
|
-
`1.7s × n`. A cold app launch is a **one-off cost amortised over the route**:
|
|
766
|
-
|
|
767
|
-
```
|
|
768
|
-
per step = (model round trip + launch cost + ~1.7s × n) / n
|
|
769
|
-
```
|
|
770
|
-
|
|
771
|
-
That reconciles a 7-step replay at 1.82 s/step with a 2-step route at 6.2 s/step
|
|
772
|
-
— one launch spread over 7 steps or over 2. **The `~1.7 s` figure is warm taps
|
|
773
|
-
inside a batch**, and short routes are materially worse than the table above
|
|
774
|
-
implies on its own.
|
|
775
|
-
|
|
776
|
-
`step_ratio` was **1** throughout: when a run completed it took exactly the
|
|
777
|
-
minimum number of steps.
|
|
778
|
-
A field report put the split at **34% simframe, 60% agent round trips** over
|
|
779
|
-
462 s of wall clock — the tester's *"30+ seconds between each step"* was
|
|
780
|
-
accurate and was not simframe. So there is nothing left to win inside the
|
|
781
|
-
engine, and the only variable is `n`: `simframe hpi` reports `steps_per_call`
|
|
782
|
-
for exactly that reason.
|
|
783
|
-
|
|
784
|
-
Which is why **every hard-fail that drops a caller back to single-stepping is a
|
|
785
|
-
latency bug**. Between two field reports on the same flow, one run took 33 tool
|
|
786
|
-
calls and the next took **16**, for 28 executed steps and no screenshots at
|
|
787
|
-
all — the difference being defects fixed, not code made faster.
|
|
788
|
-
|
|
789
|
-
## How it works
|
|
790
|
-
|
|
791
|
-
```
|
|
792
|
-
┌─── simframed — one Swift daemon per simulator ──────────────────────────────┐
|
|
793
|
-
│ │
|
|
794
|
-
│ display damage callback ──► read IOSurface ──► scale ──► hash │
|
|
795
|
-
│ the screen tells us 0.13 ms 6.6 ms total │
|
|
796
|
-
│ │ │
|
|
797
|
-
│ Vision OCR reads the same surface ─────┤ no PNG, no file, no spawn │
|
|
798
|
-
│ │ │
|
|
799
|
-
│ Indigo HID ◄── control socket ◄────────┤ 0600, one JSON object per line │
|
|
800
|
-
│ taps, swipes, text │ │
|
|
801
|
-
│ ▼ │
|
|
802
|
-
│ ~/.simframe/<udid>/ frames · state.json · meta.json │
|
|
803
|
-
└─────────────────────────────────────────────────────────────────────────────┘
|
|
804
|
-
│ a rename is atomic
|
|
805
|
-
┌────────────────────────────────────▼────────────────────────────────────────┐
|
|
806
|
-
│ MCP server / CLI: stat + read, or one socket round trip for input. │
|
|
807
|
-
│ Screen memory: layout hash ──► label → point, built once per screen. │
|
|
808
|
-
└─────────────────────────────────────────────────────────────────────────────┘
|
|
809
|
-
```
|
|
810
|
-
|
|
811
|
-
The daemon links CoreSimulator and SimulatorKit, which are private frameworks
|
|
812
|
-
with no documentation and no stability promise. Everything it calls is recorded
|
|
813
|
-
in [`docs/PRIVATE_API.md`](docs/PRIVATE_API.md) with the evidence behind it, so
|
|
814
|
-
an Xcode upgrade that moves something is a bounded fix rather than an
|
|
815
|
-
archaeology project. If a layer breaks, simframe degrades to the layer below
|
|
816
|
-
and `doctor` says which.
|
|
817
|
-
|
|
818
|
-
Run `simframe start --engine=screenshot` to use the one-frame-at-a-time loop
|
|
819
|
-
instead — it is also the only capture engine on Android, where it reaches the
|
|
820
|
-
emulator console rather than any simulator tool. `--engine=simctl` is still
|
|
821
|
-
accepted as the name that loop used to have.
|
|
822
|
-
|
|
823
|
-
- **Files are the IPC for reads.** The daemon renames completed frames into
|
|
824
|
-
place; readers just read them. A rename is atomic, so a reader can never see a
|
|
825
|
-
half-written frame. Input is the one thing that needs a reply, and it goes
|
|
826
|
-
over a `0600` Unix socket — the file system is the whole permission model.
|
|
827
|
-
- **Almost no dependencies.** The only runtime npm dependency is the MCP SDK.
|
|
828
|
-
The daemon is Swift built from source against frameworks already on the
|
|
829
|
-
machine.
|
|
830
|
-
- **The screen says when it changed.** The capture loop is driven by the
|
|
831
|
-
display's damage callback rather than a timer, so an idle screen costs
|
|
832
|
-
nothing and a moving one is picked up at once.
|
|
833
|
-
- **It backs off when nothing happens.** 4 fps while the screen moves, 1.5 fps
|
|
834
|
-
once still, snapping back instantly on change.
|
|
835
|
-
- **One writer per device.** Ownership lives in `meta.json`; `stop` refuses to
|
|
836
|
-
kill a loop another client is using unless forced.
|
|
837
|
-
- **A wedged capture loop never looks like a calm screen.** Every answer carries
|
|
838
|
-
a liveness check, and `wait` fails loudly rather than quietly timing out.
|
|
839
|
-
- **An action with no visible effect is reported, not waited out.** Selecting a
|
|
840
|
-
radio button moves ~0.1 % of the screen — below the change threshold — which
|
|
841
|
-
used to burn the full timeout. Now the step returns in ~3 s marked
|
|
842
|
-
`[no visible change]`, so you know to check rather than wait.
|
|
843
|
-
- **And a gesture aimed at a coordinate says what it landed on.** A tap or swipe
|
|
844
|
-
that changed nothing prints the element covering its start point, because a
|
|
845
|
-
gesture goes to whatever is on top there:
|
|
846
|
-
|
|
847
|
-
```
|
|
848
|
-
swiped 83,141 -> 83,800 [no visible change]
|
|
849
|
-
[the swipe start point 83,141 is inside "Settings" (nav-bar)]
|
|
850
|
-
```
|
|
851
|
-
|
|
852
|
-
Reported from the field as fifteen minutes lost to six identical swipes that
|
|
853
|
-
a support banner was swallowing — a banner that was *in the element list the
|
|
854
|
-
same call printed*. Nothing said "there is something at y≈753 and you started
|
|
855
|
-
at y=750".
|
|
856
|
-
- **A control half off the edge is offered, clamped, and says so.** A filter
|
|
857
|
-
chip with 29 pt of itself visible used to be dropped — the off-screen test
|
|
858
|
-
reads the *centre* — and OCR's reading of that same sliver, a box labelled
|
|
859
|
-
`Flc`, was printed in its place. So the map did not merely omit a control; it
|
|
860
|
-
offered a meaningless name for it at an ordinary-looking coordinate. Now:
|
|
861
|
-
|
|
862
|
-
```
|
|
863
|
-
#6 button 387,191 Flowering (partly off-screen — the coordinate is the middle of the visible part)
|
|
864
|
-
```
|
|
865
|
-
|
|
866
|
-
`scrollTo` is deliberately unchanged: a clipped element is still "not in
|
|
867
|
-
view", so it still scrolls rather than calling a sliver good enough.
|
|
868
|
-
|
|
869
|
-
## CLI
|
|
870
|
-
|
|
871
|
-
The CLI is the low-token path, and it is a first-class one: `--json` is on every
|
|
872
|
-
command, so nothing has to be parsed out of prose.
|
|
873
|
-
|
|
874
|
-
```bash
|
|
875
|
-
simframe ui # the numbered screen map — start here
|
|
876
|
-
simframe ui --json | jq '.elements[] | select(.type=="button") | .label'
|
|
877
|
-
simframe do flow.json # a whole flow, verified, then the end-state map
|
|
878
|
-
simframe do flow.json --save=checkout # save it if every step verified
|
|
879
|
-
simframe flow run checkout # replay it
|
|
880
|
-
simframe tap "#3" # or "Save", or "@120,400"
|
|
881
|
-
simframe find "the save button" # resolve an intent without acting
|
|
882
|
-
simframe screens # screens this device has learned
|
|
883
|
-
simframe goto invoices # walk to a known screen over known steps
|
|
884
|
-
simframe state --since=$H # what changed, as text
|
|
885
|
-
simframe mark # hash of the current frame, for --since
|
|
886
|
-
simframe wait --since=$H # change, then settle
|
|
887
|
-
simframe recall # what happened in the last minute, as text
|
|
888
|
-
simframe recall --ago=15000 # the frame from 15s ago
|
|
889
|
-
simframe frame --out=now.png # newest frame, native resolution, to a file
|
|
890
|
-
simframe strip --count=6 # contact sheet, for an animation
|
|
891
|
-
simframe doctor --strict # any degraded layer is a non-zero exit
|
|
892
|
-
simframe escalations # why simframe still needs a model, by reason
|
|
893
|
-
simframe escalations --session # ...this agent only, not every agent on the device
|
|
894
|
-
simframe supervisions # local supervisor rulings, and what came of each
|
|
895
|
-
simframe hpi # speed and accuracy against a human baseline
|
|
896
|
-
simframe baseline record settings-larger-text --runs=5 # record the human
|
|
897
|
-
simframe input reset # rebuild the HID session, without restarting anything
|
|
898
|
-
simframe diagnose # what this device is doing right now, and which failure it is
|
|
899
|
-
simframe revive # power-cycle a wedged device: stop, shutdown, boot, start, reset input
|
|
900
|
-
simframe start / status / stop [--force] / devices
|
|
901
|
-
simframe ui --device=emulator-5554 # or export SIMFRAME_DEVICE once
|
|
902
|
-
```
|
|
903
|
-
|
|
904
|
-
`--device` takes `--device=X` and `--device X` alike. It used to take only the
|
|
905
|
-
first: the space form set the flag to `true` and then resolved a device named
|
|
906
|
-
"true", which is a poor answer to a flag `doctor`'s own advice tells you to
|
|
907
|
-
type.
|
|
908
|
-
|
|
909
|
-
### When the simulator stops rendering
|
|
910
|
-
|
|
911
|
-
A simulator driven hard for several minutes can stop rendering: capture fails,
|
|
912
|
-
`simctl screenshot` fails too, and the daemon's own recoveries — re-resolving the
|
|
913
|
-
display port, then rebinding the device — do not help. It reports the state
|
|
914
|
-
rather than acting on it, because a capture loop that rebooted the device it was
|
|
915
|
-
watching would be a tool reaching for the mains:
|
|
916
|
-
|
|
917
|
-
```
|
|
918
|
-
capture is wedged and both recoveries are spent (2 port re-resolves, 2 device
|
|
919
|
-
rebinds, no frame since). This needs the device restarted —
|
|
920
|
-
`simframe revive --device=<udid>`. Backing off until a frame arrives.
|
|
921
|
-
```
|
|
922
|
-
|
|
923
|
-
**`simframe diagnose` says which failure it is**, which `doctor` cannot: `doctor`
|
|
924
|
-
answers "can this machine capture", and this answers "what is this device doing
|
|
925
|
-
right now". It reads only what discriminates — frame sequence, age and
|
|
926
|
-
stillness, element counts split by sensor, how many of them fuse, and who holds
|
|
927
|
-
the front by pid — and returns one of:
|
|
928
|
-
|
|
929
|
-
| verdict | what it means |
|
|
930
|
-
| --- | --- |
|
|
931
|
-
| `capture-down` | no frames at all, carrying the daemon's own sentence |
|
|
932
|
-
| `nothing-readable` | frames arriving, neither sensor finds a single element |
|
|
933
|
-
| `stale-frame` | both sensors full, almost nothing fuses — the framebuffer is behind the tree, so **an image from this device is not safe to trust** |
|
|
934
|
-
| `not-presenting` | an app holds the front by pid and the display shows almost nothing |
|
|
935
|
-
| `healthy` | — |
|
|
936
|
-
|
|
937
|
-
`not-presenting` deliberately does **not** say whether that is a lock screen, a
|
|
938
|
-
dead surface or a crashed system shell. It is not knowable from here, and
|
|
939
|
-
guessing is how a regex ended up standing where a measurement belongs.
|
|
940
|
-
|
|
941
|
-
The `stale-frame` threshold is measured rather than chosen: element fusion on
|
|
942
|
-
five healthy screens ran 0.667–0.929, so the threshold sits at 0.1 — 6.7× below
|
|
943
|
-
the observed floor rather than inside the metric's own noise. Numbers in
|
|
944
|
-
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
945
|
-
|
|
946
|
-
`simframe revive` is that restart, in the order that matters — stop the daemon,
|
|
947
|
-
shut the device down, boot it and *wait for the boot to finish*, start capture,
|
|
948
|
-
rebuild the HID session — and it ends by checking frames are flowing again
|
|
949
|
-
rather than by reporting that the steps ran. It is a command and not a
|
|
950
|
-
behaviour: the decision stays yours.
|
|
951
|
-
|
|
952
|
-
### Reading what the local supervisor decided
|
|
953
|
-
|
|
954
|
-
With `SIMFRAME_SUPERVISOR=apple`, every consultation is written to
|
|
955
|
-
`~/.simframe/<udid>/supervisions.jsonl` with what the executor observed
|
|
956
|
-
afterwards, which is what makes a ruling scoreable rather than merely recorded.
|
|
957
|
-
`simframe supervisions` reads it:
|
|
958
|
-
|
|
959
|
-
```
|
|
960
|
-
14 supervisor rulings on iPhone 17 Pro
|
|
961
|
-
|
|
962
|
-
stop -> stopped 7
|
|
963
|
-
wait -> recovered 5
|
|
964
|
-
wait -> still_failed 2
|
|
965
|
-
|
|
966
|
-
sourced: model 14
|
|
967
|
-
median latency: 1440ms
|
|
968
|
-
edges the graph had timed: 0/14 — 14 ruling(s) are on edges with no p95
|
|
969
|
-
```
|
|
970
|
-
|
|
971
|
-
That last line is the honest one: a step that failed is usually a step that has
|
|
972
|
-
never succeeded on that edge, so the graph has no timing to compare against.
|
|
973
|
-
|
|
974
|
-
### Keeping a session cheap
|
|
975
|
-
|
|
976
|
-
The expensive part of driving a simulator with an agent is not the tapping, it
|
|
977
|
-
is the thinking between taps — observe, think, tap, observe, think. Measured
|
|
978
|
-
over one real session against a third-party app: **62 tool calls for 179 steps**,
|
|
979
|
-
and 48 of those calls were three steps or fewer. A twelve-step flow arrived as
|
|
980
|
-
five calls, and every boundary between them was a think.
|
|
981
|
-
|
|
982
|
-
Three things move that number, and simframe does the first two for you:
|
|
983
|
-
|
|
984
|
-
- **Batch.** `sim_do` runs a whole flow in one call, with an assert after each
|
|
985
|
-
step that matters. The asserts are what make it safe not to look in between:
|
|
986
|
-
a step that lands somewhere unplanned halts the flow instead of letting the
|
|
987
|
-
next four run against the wrong screen.
|
|
988
|
-
- **A `next:` line on every action result** — from the CLI and the MCP server
|
|
989
|
-
alike — computed locally from what the daemon already knows — whether the screen settled, whether the graph
|
|
990
|
-
recognises it, how many elements it has, whether any labels repeat. When it
|
|
991
|
-
says *nothing ambiguous — chain the next steps in one sim_do without looking
|
|
992
|
-
again*, that is the tool telling the agent it does not need to think.
|
|
993
|
-
- **A trailing map that was re-read, not recalled.** An action pays one
|
|
994
|
-
perception pass — a few hundred milliseconds, locally — so the map it returns
|
|
995
|
-
is the screen as it is now. The alternative was an agent spending a whole turn
|
|
996
|
-
on `ui --refresh` because it could not trust the one it was given.
|
|
997
|
-
- **One goal per session.** Sessions get slower with every turn. A flow that
|
|
998
|
-
runs as one call adds one exchange to the context instead of twelve.
|
|
999
|
-
|
|
1000
|
-
And one thing to know about waiting: `settle` asks whether the screen stopped
|
|
1001
|
-
moving, and a screen waiting on a network call has stopped moving. For anything
|
|
1002
|
-
that arrives over the network, assert on the content you expect —
|
|
1003
|
-
`{"waitFor": {"value": "Kate Bell"}}` — rather than on stillness. The map says
|
|
1004
|
-
`STILL LOADING` when the classifier can see a load in flight, but only you know
|
|
1005
|
-
what "arrived" means.
|
|
1006
|
-
|
|
1007
|
-
And the expensive habit worth naming: in that session, **28 of 62 calls returned
|
|
1008
|
-
a screenshot** — about a third of its entire token cost — because the text map
|
|
1009
|
-
could not report what a text field contained. It can now, so check the map
|
|
1010
|
-
before reaching for pixels: a row carries the element's contents (`= Fryer 3`)
|
|
1011
|
-
and its state (`disabled`), and the flow's own verdict already said whether the
|
|
1012
|
-
action worked.
|
|
1013
|
-
|
|
1014
|
-
### The Claude Code skill
|
|
1015
|
-
|
|
1016
|
-
[`skills/simframe/SKILL.md`](skills/simframe/SKILL.md) teaches the CLI path
|
|
1017
|
-
directly: the cheap-to-expensive order, the selector grammar, what each verdict
|
|
1018
|
-
means and what to do about it. It ships with the package, so an installed copy
|
|
1019
|
-
has it.
|
|
1020
|
-
|
|
1021
|
-
```bash
|
|
1022
|
-
mkdir -p ~/.claude/skills
|
|
1023
|
-
ln -s "$(npm root -g)/simframe/skills/simframe" ~/.claude/skills/simframe
|
|
1024
|
-
```
|
|
1025
|
-
|
|
1026
|
-
A skill and an MCP server are not redundant. The MCP server is discoverable —
|
|
1027
|
-
it appears in the tool list without anybody setting it up. The skill is cheaper:
|
|
1028
|
-
Claude reads 3–5 lines of CLI output instead of a tool result, and none of the
|
|
1029
|
-
MCP schema is in context until a tool is actually used. Ship both, use whichever
|
|
1030
|
-
the client makes easy.
|
|
1031
|
-
|
|
1032
|
-
## Degrading is allowed. Degrading quietly is not
|
|
1033
|
-
|
|
1034
|
-
simframe is built to degrade rather than fail: no Swift toolchain still gives
|
|
1035
|
-
you frames through `simctl`, no accessibility tree still gives you OCR. That
|
|
1036
|
-
policy is right, and it nearly sank the tool twice — because a downgrade looked
|
|
1037
|
-
exactly like everything working.
|
|
1038
|
-
|
|
1039
|
-
Once, one file was missing from the published package, so `Package.swift`
|
|
1040
|
-
declared a test target with no directory, SwiftPM reported overlapping sources,
|
|
1041
|
-
and **every install silently fell back to the slow engine**. Another time OCR
|
|
1042
|
-
shipped disabled the same way. Both passed the tests. Both printed nothing. The
|
|
1043
|
-
bug was never the missing file; it was the silence.
|
|
1044
|
-
|
|
1045
|
-
So every downgrade now announces itself:
|
|
1046
|
-
|
|
1047
|
-
- `simframe start` prints the engine it chose, and if it is the slow one, why —
|
|
1048
|
-
build error, missing sources, or "reason unrecorded" if the daemon was started
|
|
1049
|
-
by an earlier process.
|
|
1050
|
-
- `simframe doctor` marks a degraded layer `WARN`, not `ok`, and summarises what
|
|
1051
|
-
is degraded and what that costs.
|
|
1052
|
-
- A dependency that is simply not installed is `--`, not `WARN`. The distinction
|
|
1053
|
-
is deliberate: `WARN` means this machine could be doing better and silently is
|
|
1054
|
-
not, which is the failure worth shouting about. An optional fallback missing on
|
|
1055
|
-
a fresh machine has not degraded from anything, and `--strict` ignores it.
|
|
1056
|
-
- A device whose capture has **wedged** says so: `capture: stalled — the display
|
|
1057
|
-
surface has been unreadable for 62s; 3 re-attaches did not help; only
|
|
1058
|
-
restarting the device is known to cure it`. This is a different thing from a
|
|
1059
|
-
still screen, and it used to look identical, because a damage-driven engine
|
|
1060
|
-
produces no frames for either. An agent told "nothing changed" keeps tapping;
|
|
1061
|
-
one told the simulator is wedged stops. simframe reports it and does not
|
|
1062
|
-
restart your device.
|
|
1063
|
-
- `--strict`, or `SIMFRAME_STRICT=1`, turns any downgrade into a non-zero exit.
|
|
1064
|
-
CI runs strict, so a release cannot ship in the state that shipped twice.
|
|
1065
|
-
|
|
1066
|
-
```
|
|
1067
|
-
$ simframe doctor
|
|
1068
|
-
ok capture engine simframed
|
|
1069
|
-
WARN input driver idb — the daemon's control socket is not up
|
|
1070
|
-
WARN accessibility tree idb — the host-side translator did not load
|
|
1071
|
-
```
|
|
1072
|
-
|
|
1073
|
-
Two checks enforce it. A packaging check derives the required file list from the
|
|
1074
|
-
build's own inputs — a hand-written list is what rotted last time — and runs in
|
|
1075
|
-
seconds without a simulator. An integration job installs the packed tarball on a
|
|
1076
|
-
real simulator and asserts `capture.engine`, `input.driver` and `ocr.available`
|
|
1077
|
-
are all the good values, under `--strict`.
|
|
172
|
+
| [`docs/GUIDE.md`](docs/GUIDE.md) | the full guide: recovering without a round trip, screen memory and the two hashes, navigating by memory, the CLI, diagnosing a wedged simulator, keeping a session cheap, the roadmap |
|
|
173
|
+
| [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | every number, with N, median, p95, machine and the measurement traps |
|
|
174
|
+
| [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
|
|
175
|
+
| [`docs/DECISIONS.md`](docs/DECISIONS.md) | the judgements that changed the plan, including a phase cancelled by its own measurement |
|
|
176
|
+
| [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
|
|
177
|
+
| [`docs/ARTICLE.md`](docs/ARTICLE.md) | *Agents shouldn't blink* — the argument the whole thing adds up to |
|
|
1078
178
|
|
|
1079
|
-
|
|
1080
|
-
|
|
1081
|
-
|
|
1082
|
-
|
|
179
|
+
Almost everything here was decided by a measurement rather than an argument,
|
|
180
|
+
and several measurements reversed a decision that had already shipped. Two
|
|
181
|
+
headline results: a warm transition graph removes about 27% of an agent's
|
|
182
|
+
model round trips, and a one-line threshold on a number the daemon already
|
|
183
|
+
computes beat every local model tried as a step supervisor — 95% against
|
|
184
|
+
77–91%, at zero latency.
|
|
1083
185
|
|
|
1084
186
|
## Limitations
|
|
1085
187
|
|
|
1086
|
-
- Simulators and
|
|
1087
|
-
|
|
1088
|
-
|
|
1089
|
-
|
|
1090
|
-
|
|
1091
|
-
|
|
1092
|
-
|
|
1093
|
-
|
|
1094
|
-
|
|
1095
|
-
-
|
|
1096
|
-
|
|
1097
|
-
|
|
1098
|
-
|
|
1099
|
-
|
|
1100
|
-
`sim_do` does by default.
|
|
1101
|
-
- Region maps need a baseline inside the ~90 s history window. Older baselines
|
|
1102
|
-
still get a reliable changed / did-not-change, without a map of what moved.
|
|
1103
|
-
- Screen memory assumes a screen's layout is stable. A screen that reflows
|
|
1104
|
-
dramatically between visits will simply be rebuilt.
|
|
1105
|
-
- It speeds up *confirming* a fix, not *locating* one. A bug living in a memo
|
|
1106
|
-
comparator or a stale closure is not visible in any frame.
|
|
1107
|
-
- A switch is tapped at its **activation point** when the app publishes one —
|
|
1108
|
-
UIKit's `accessibilityActivationPoint`, which is what a switch answers with,
|
|
1109
|
-
and tapping it flipped a real switch 3 of 3 times where the frame centre
|
|
1110
|
-
managed 0 of 3. Only **4 of 75** elements on a measured screen publish one,
|
|
1111
|
-
though, so where the app says nothing the tap still goes to the centre of the
|
|
1112
|
-
frame — and a switch's frame is the whole row, so it lands on the label and
|
|
1113
|
-
the control at the trailing end does not move. `@x,y` remains the escape
|
|
1114
|
-
hatch there. Never a guessed offset: nil means the app did not answer.
|
|
1115
|
-
- A launch is confirmed to have fronted **on iOS only**. It compares the pid
|
|
1116
|
-
`simctl launch` printed against the pid the device reports as frontmost, in
|
|
1117
|
-
2–5 ms. Android reports neither, so a launch there is not checked — `am start`
|
|
1118
|
-
fronts synchronously, which is why it has not bitten, but that is not a check
|
|
1119
|
-
and `doctor` says so. Note that a launch which starts a process without
|
|
1120
|
-
bringing it forward now **fails** rather than returning success: a flow that
|
|
1121
|
-
used to pass through such a launch and then assert on the previous app's
|
|
1122
|
-
screen will start failing, correctly.
|
|
1123
|
-
- The simulator's display pipeline stops rendering under rapid app relaunch —
|
|
1124
|
-
about six cycles, reproducibly — and every frame comes back black while
|
|
1125
|
-
`simctl` itself reports success. simframe now says so instead of reading a
|
|
1126
|
-
black screen as a calm one, but it cannot fix it: restarting the device is
|
|
1127
|
-
the cure that always works, and it usually recovers on its own.
|
|
1128
|
-
|
|
1129
|
-
## What was decided, and what was not built
|
|
1130
|
-
|
|
1131
|
-
[`docs/DECISIONS.md`](docs/DECISIONS.md) is the register of judgements that
|
|
1132
|
-
changed the plan: a phase cancelled by its own measurement, two features built
|
|
1133
|
-
and reverted for cause, and the premises that turned out to be false. It is
|
|
1134
|
-
short on purpose — the numbers live in
|
|
1135
|
-
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) and the open work in
|
|
1136
|
-
[`docs/DEFERRED.md`](docs/DEFERRED.md).
|
|
1137
|
-
|
|
1138
|
-
The most useful entry is a **no-go**: Phase 17 proposed a small on-device model
|
|
1139
|
-
to choose the next element, and measuring the prize before the model showed the
|
|
1140
|
-
matcher already resolves 37 of 40 real decisions. By the time a step reaches
|
|
1141
|
-
simframe the decision has already been made — the goal names the option,
|
|
1142
|
-
because the agent chose it and then asked for it by name.
|
|
1143
|
-
|
|
1144
|
-
## Roadmap
|
|
1145
|
-
|
|
1146
|
-
- **Extend the confirm vocabulary beyond English.**
|
|
1147
|
-
- **Region bands from clustering**, replacing the positional bands. They have
|
|
1148
|
-
produced three bugs in three phases, and on Android they put a URL bar in the
|
|
1149
|
-
nav bar and its URL into the screen's identity.
|
|
1150
|
-
- **Phase 8b, conditionally:** an instrumentation APK for the Android
|
|
1151
|
-
accessibility tree, with the criteria for doing it stated in
|
|
1152
|
-
`docs/DEFERRED.md` rather than left to enthusiasm.
|
|
188
|
+
- Simulators and emulators only; nothing here can reach a physical device.
|
|
189
|
+
- The daemon links private frameworks. They survived the iOS 26 transition,
|
|
190
|
+
but an Xcode upgrade can move a symbol; `doctor` names the broken layer and
|
|
191
|
+
`--engine=screenshot` still works.
|
|
192
|
+
- Android has no accessibility tree, so screen identity there rests on OCR
|
|
193
|
+
geometry alone: thinner and noisier than iOS.
|
|
194
|
+
- Hardware buttons: only `home`. The other codes are unverified and a wrong
|
|
195
|
+
one can crash the device, so they return an error rather than a guess.
|
|
196
|
+
- The confirm vocabulary (`OK`, `Save`, `Done`…) is English.
|
|
197
|
+
- It speeds up *confirming* a fix, not *locating* one. A bug in a stale
|
|
198
|
+
closure is not visible in any frame.
|
|
199
|
+
|
|
200
|
+
The full list, with the measurements behind each item, is in
|
|
201
|
+
[the guide](docs/GUIDE.md#limitations).
|
|
1153
202
|
|
|
1154
203
|
## Releasing
|
|
1155
204
|
|
|
1156
|
-
`npm version`
|
|
1157
|
-
|
|
1158
|
-
|
|
1159
|
-
|
|
1160
|
-
|
|
1161
|
-
git push --follow-tags
|
|
1162
|
-
```
|
|
1163
|
-
|
|
1164
|
-
Before that hook existed, `server.json` had to be hand-edited between two
|
|
1165
|
-
commands, and the release that forgot failed at the workflow's own agreement
|
|
1166
|
-
check — which is the one thing that check is for.
|
|
1167
|
-
|
|
1168
|
-
The `release` workflow verifies tag/`package.json`/`server.json` agree, validates
|
|
1169
|
-
`server.json` against the live registry, and publishes to npm and the MCP
|
|
1170
|
-
Registry. It holds **no secrets** — both halves authenticate with the workflow's
|
|
1171
|
-
GitHub OIDC identity.
|
|
1172
|
-
|
|
1173
|
-
That needs one setup step on npmjs.com, not in this repo: the package must have a
|
|
1174
|
-
Trusted Publisher pointing at this repository and `release.yml` (Package →
|
|
1175
|
-
Settings → Trusted Publisher → GitHub Actions). Without it npm has nothing to
|
|
1176
|
-
trust and fails with `ENEEDAUTH`. npm is ending token publishing in January 2027,
|
|
1177
|
-
and the tokens that work in CI need 2FA bypass, which npm's own UI warns against —
|
|
1178
|
-
so OIDC is the durable path, not merely the tidier one.
|
|
205
|
+
`npm version` is the only way to bump a version — its hook keeps `server.json`
|
|
206
|
+
in step — and `git push --follow-tags` runs a workflow that publishes to npm
|
|
207
|
+
and the official MCP Registry with no secrets. Why it is built that way, and
|
|
208
|
+
the release that once vanished for a week, is in
|
|
209
|
+
[the guide](docs/GUIDE.md#releasing).
|
|
1179
210
|
|
|
1180
211
|
## License
|
|
1181
212
|
|