simframe 0.18.0 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,152 +3,27 @@
3
3
  [![ci](https://github.com/lvlrSajjad/simframe/actions/workflows/ci.yml/badge.svg)](https://github.com/lvlrSajjad/simframe/actions/workflows/ci.yml)
4
4
  [![npm](https://img.shields.io/npm/v/simframe.svg)](https://www.npmjs.com/package/simframe)
5
5
  [![license](https://img.shields.io/npm/l/simframe.svg)](./LICENSE)
6
+ [![Glama score](https://glama.ai/mcp/servers/lvlrSajjad/simframe/badges/score.svg)](https://glama.ai/mcp/servers/lvlrSajjad/simframe)
6
7
 
7
- **Eyes, hands and memory for an agent driving the iOS Simulator or an Android emulator.**
8
+ **Eyes, hands and memory for a coding agent driving the iOS Simulator or an
9
+ Android emulator.** The agent reads the screen as text with tap points, runs a
10
+ whole flow in one call, and remembers screens it has seen — so it stops paying
11
+ a screenshot, a model round trip and a re-read for every single step.
8
12
 
9
- [Website](https://lvlrsajjad.github.io/simframe/) · [npm](https://www.npmjs.com/package/simframe)
13
+ ![simframe driving Settings on an iPhone 17 Pro simulator: one call, four verified steps, no screenshots, and the screen returned as text with tap points](https://raw.githubusercontent.com/lvlrSajjad/simframe/main/docs/simframe-demo.gif)
10
14
 
11
- An agent driving a simulator is slow for three reasons, and only the first
12
- one is obvious:
15
+ ## Install in ten seconds
13
16
 
14
- 1. **Every look is a wait.** `simctl io screenshot` costs ~130 ms of blocking
15
- latency, paid again on every glance — and paid twice whenever the agent
16
- captures mid-animation and has to look again.
17
- 2. **Every step is a round trip.** Tap, screenshot, reason, tap, screenshot. A
18
- twelve-step flow costs twelve model turns, and the model turns cost far more
19
- than the milliseconds.
20
- 3. **Nothing is remembered.** The same screen gets re-read and re-reasoned about
21
- every single time it appears.
17
+ You need a Mac with Xcode (you have one if you have a simulator) and Node 18+.
18
+ Nothing else — no idb, no Appium, no Python.
22
19
 
23
- And there is a fourth that is pure waste: **an image is the most expensive way
24
- to ask what is on screen.** A screenshot costs ~1,600 tokens when it is handled
25
- as a native image block and 15,000–25,000 when it is not, and it does not tell
26
- you what is tappable or where — you have to measure that by eye.
27
-
28
- simframe attacks all four: a background loop keeps the newest frame warm, whole
29
- flows run in one call, screens the agent has seen before are answered from
30
- memory, and every answer is text with tap points in it. Nothing returns an
31
- image unless you ask for one.
32
-
33
- ### It is also a study case, and the numbers are the point
34
-
35
- Almost everything here was decided by a measurement rather than by an argument,
36
- and several of those measurements **reversed a decision that had already
37
- shipped**. They are written down in full, with what we expected beforehand, what
38
- it cost to find out, and the mistakes made getting there — because a benchmark
39
- that only records the winner teaches nothing.
40
-
41
- | | |
42
- | --- | --- |
43
- | [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | the record: N, median, p95, machine, every round |
44
- | [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
45
- | [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
46
- | [`docs/ARTICLE.md`](docs/ARTICLE.md) | the argument the whole thing adds up to |
47
-
48
- The headline results, all free to reuse:
49
-
50
- - **A warm transition graph removes ~27% of an agent's model round trips** —
51
- isolated with a three-pass field experiment, and corroborated to three points
52
- by an independent estimate. Operator knowledge is worth twice that, and no
53
- tool here carries it yet. ([§14](docs/EXPERIMENTS.md))
54
- - **A bigger local judge is not a better one.** `qwen3:14b` scored *lower* than
55
- `qwen3:8b` — 82% against 91% — while being 79% larger and 62% slower. A
56
- one-line threshold on a number already computed beat all of them at **95% and
57
- zero latency**.
58
- - **Every interactive element came from the accessibility tree — 72 of 72 — and
59
- 83% of them have no text at all.** Screen *recognition*, though, did not need
60
- the tree once. The ladder splits there: semantics need a tree, identity needs
61
- pixels.
62
- - **Capture is ~31× faster than a screenshot** for the same work, and the
63
- capture primitive itself about a thousand times faster.
64
- - **Three independent reporters converged on one thing**: a refusal is cheap and
65
- a confident wrong answer is expensive. Every serious bug they found was a
66
- component reporting more certainty than it had.
67
-
68
- ## What changed, measured
69
-
70
- Same four-tab navigation flow, on a real production app:
71
-
72
- | | Before | With simframe |
73
- | --- | --- | --- |
74
- | Look at the screen | ~130–400 ms, blocking | **~20 ms**, already captured |
75
-
76
- | "Did anything change?" | a full image | **~2 ms**, text only |
77
- | Finding a control | read tree (~570 ms) + reason | **~1 ms** from memory |
78
- | A 4-step flow, verified | 4+ model round trips | **1 call**, 3.6 s |
79
- | Same flow, 3rd run | no improvement — every run is the first | **3.7 s, 4/4 from memory, 4/4 verified** |
80
- | A 10-step flow | 10 turns, 10 images (~16,000 tokens at best) | **1 turn, 0 images, ~1,650 characters** |
81
- | Reading a screen | an image: ~1,600 tokens, no tap points | **~330 tokens** of text, with tap points |
82
-
83
- Every figure above is the cost inside a live process — the MCP server, or the
84
- daemon answering a socket — which is how an agent actually uses simframe. A
85
- one-shot `simframe` command from a shell pays about 200 ms of Node startup on
86
- top, and a frame sitting on an idle screen can be older than 20 ms because the
87
- capture loop throttles when nothing moves. `~20 ms` is the read, not the
88
- process.
89
-
90
- The four-tab tour, three times back to back from a cleared memory:
91
-
92
- | Pass | Wall clock | Steps verified | Controls from memory |
93
- | --- | --- | --- | --- |
94
- | 1 | 10.2 s | 0/4 — nothing is known yet | **4/4** |
95
- | 2 | **3.6 s** | **4/4** | **4/4** |
96
- | 3 | **3.7 s** | **4/4** | **4/4** |
97
-
98
- Every step is checked against what the same action did last time, and the run
99
- records its own preconditions — which input path, which daemon, whether the
100
- daemon was replaced mid-run — so a regression shows up in the measurement rather
101
- than hiding inside it. Earlier versions of this table quoted 7.4 s → 3.3 s with
102
- verification switched off; those numbers were measured while input was silently
103
- falling back to a slower path and the capture daemon was being replaced by every
104
- command, so they measured two bugs rather than the tool.
105
-
106
- ## Install
107
-
108
- ```bash
109
- npm install -g simframe
110
- simframe doctor
111
- ```
112
-
113
- **Whichever of those two commands you run first** builds a small Swift daemon
114
- from source — including `doctor`, which is why a cold `doctor` takes around 15
115
- seconds and every later one takes two. It needs the Xcode command line tools,
116
- which you already have if you have a simulator. Without them simframe falls
117
- back to the original `simctl` loop and says so.
118
-
119
- If more than one simulator is booted, name the one you mean — `--device=<udid>`,
120
- or `export SIMFRAME_DEVICE=<udid>` once per shell. simframe refuses to choose
121
- for you, because the first booted device is nobody's idea of "yours" and the
122
- command that would act on it is a tap.
123
-
124
- `doctor` checks each capability separately and tells you what you have:
125
-
126
- ```
127
- ok xcrun xcrun version 72.
128
- ok sips available
129
- ok input driver simframed: Indigo HID
130
- ok accessibility tree simframed: AXPTranslator, host-side
131
- ok on-device OCR available
132
- ok booted simulator iPhone 17 Pro (iOS 26.5)
133
- ok capture frame #888 322x700 in 2ms (age 538ms)
134
- ok sensor mode full — accessibility and OCR fused on every read (~164ms)
135
- ok local supervisor none — not requested (SIMFRAME_SUPERVISOR is unset)
136
- ok local planner none — not requested (SIMFRAME_PLANNER is unset)
137
- ```
138
-
139
- The last three are experiments and `none` is their normal answer. See
140
- [Local tiers, off by default](#local-tiers-off-by-default).
141
-
142
- ### Claude Code
20
+ **Claude Code**
143
21
 
144
22
  ```bash
145
23
  claude mcp add --scope user simframe -- npx -y simframe mcp
146
24
  ```
147
25
 
148
- `--scope user` makes it available in every session; without it the server is
149
- registered only for the directory you ran the command in.
150
-
151
- ### Any other MCP client
26
+ **Claude Desktop** — Settings → Developer → Edit Config, then add:
152
27
 
153
28
  ```json
154
29
  {
@@ -158,271 +33,52 @@ registered only for the directory you ran the command in.
158
33
  }
159
34
  ```
160
35
 
161
- ## Recovering without a round trip
162
-
163
- The measured cost of driving an app is not perception — warm, an
164
- accessibility-only read is 85 ms and a fused read 142 ms. It is **round trips**:
165
- in one instrumented run, 75% of the wall time was the agent thinking and the
166
- call boundary, not simframe working. So the tools that matter most are the ones
167
- that let a batch survive a problem instead of handing it back.
168
-
169
- **Fallback selectors.** `{"tap": "Save", "or": ["Done", "Confirm"]}` — tried
170
- locally in order, only an exhausted list reaching the model. Eligible after a
171
- selector that did not *resolve* and nothing else, because retrying from a screen
172
- you did not expect to be on is a second guess. A destructive-looking label is
173
- refused as a substitute even if you list it.
174
-
175
- **Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
176
- "optional": true}` is skipped when nothing matches and runs normally when
177
- something does. It exists because the opposite cost real time: a batch that
178
- included a dismissal for a first-launch sheet lost **six correct steps** on the
179
- next run because the sheet *did not appear*. A nag screen, a permission prompt,
180
- a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
181
- is the expensive outcome — a call per step instead of a call per flow. Only a
182
- selector that resolved to **nothing** is absorbed: a target that is on screen
183
- twice is ambiguous, not absent, and must still verify. Skipped steps are
184
- reported, because "the sheet was gone" and "the sheet was dismissed" are
185
- different facts.
186
-
187
- **`{"seek": "change username", "budget": 6}`** opens containers, checks, and
188
- comes back, depth first, inside a hard budget. It **acts** — opening a door
189
- changes state — and it refuses to open anything that commits, abandons or
190
- answers. It does not tap the target; it leaves you on the screen where the target
191
- resolves.
192
-
193
- **`{"sweep": "all", "fill": {…}}`** reads a long screen a viewport at a time and
194
- fills each field while it is on screen. A form taller than the screen is only
195
- knowable in pieces — the tree publishes what is rendered — and one scroll gesture
196
- travels a non-deterministic distance, so finding a field and scrolling back to it
197
- does not work. Sweeping does: on a real web form it filled every field in one
198
- call. It detects both ends by measuring how far the *content* moved, ignoring
199
- fixed chrome, which is the only reliable signal available since nothing reports a
200
- scroll offset.
201
-
202
- **`worked here before:`** puts the graph's own vocabulary in the map, most-used
203
- first, rather than reporting a count. When the remembered controls are *not* on
204
- the screen it says so instead, because that means two screens share one
205
- fingerprint — and confident advice on a misidentified screen is how a remembered
206
- label ends up pointing at a submit button.
207
-
208
- ## Local tiers, off by default
209
-
210
- Two on-device model experiments, both `none` unless asked for, both degrading to
211
- the existing matcher-then-model ladder, and CI runs with both off. They ship no
212
- weights: Apple's Foundation Models framework has nothing to download, which is
213
- the whole reason it clears this project's non-goal on shipping model weights.
214
-
215
- | flag | what it does |
216
- |---|---|
217
- | `--sensor=ax-first` | read the accessibility tree alone (~50 ms) and pay for OCR only when a resolve fails |
218
- | `--planner=apple` | order the containers `seek` opens; it cannot choose an action |
219
- | `--supervisor=apple` | when a step fails, answer `wait`, `retry` or `stop` — nothing else — before the failure reaches the model |
220
-
221
- All three are also per-call arguments on every MCP tool, because an MCP server's
222
- environment is fixed when it spawns and comparing two modes inside one session
223
- was otherwise impossible.
224
-
225
- **What is measured and what is not.** The ranker: 5 of 6 top-1 on hand-written
226
- cases, median 564 ms warm, and on a real exploration it went to the right region
227
- in two steps where reading order wandered into version strings. `ax-first` made
228
- no measurable difference to how an agent drove a real app, with one small
229
- regression and one small win.
230
-
231
- **The supervisor, on 22 labelled failures from a seeded React Native app** —
232
- and the number to judge it by is round trips, not accuracy:
233
-
234
- | | |
235
- |---|---|
236
- | handled locally, no round trip | **9 of 22** |
237
- | escalated to Claude | 13 of 22 |
238
- | a model round trip, measured in the field | 10–16 s |
239
-
240
- So roughly 90–145 s saved on that population. Accuracy was 77 / 82 / 86% across
241
- three runs of the *same* questions — it is **not deterministic**, and every
242
- earlier single-run figure in this project carried that spread without reporting
243
- it. All of its errors were `wait` where `stop` was right, which is the cheap
244
- direction: a wrong `wait` costs a settle and a re-run, and `stop` already means
245
- "hand back to Claude with the unattempted steps" rather than "give up".
246
-
247
- **Still a bench, not the field.** These are fixtures we designed, on a testbed we
248
- built, labelled by the person who then scored them. Numbers and conditions are in
249
- [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md); the judgements, including a phase
250
- cancelled by its own measurement, are in [`docs/DECISIONS.md`](docs/DECISIONS.md).
251
-
252
- **And the measurements that changed our minds** are in
253
- [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md), with what we expected beforehand
254
- written down beside each — including the capacity comparison the supervisor
255
- design was assumed to make unnecessary. Asked the same 22 situations, three
256
- times each:
257
-
258
- | arm | accuracy | median | deterministic |
259
- |---|---|---|---|
260
- | always the commonest answer | 55% | — | — |
261
- | **`stillMs > 3000ms`, no model at all** | **95%** | **0 ms** | yes |
262
- | Apple Foundation Models (~3B, on-device) | 77 / 82 / 86% | ~640 ms | **no** |
263
- | `qwen3:8b` via Ollama (4-bit, 5.2 GB) | **91%** | 919 ms | yes |
264
- | `qwen3:14b` via Ollama (4-bit, 9.3 GB) | 82% | 1,489 ms | yes |
265
-
266
- Three results, and the third is the one that matters. The larger model scored
267
- *lower* than the smaller one. A free threshold on a number the daemon already
268
- computes beat all three, on a population it was never fitted to. And **the
269
- accuracy ranking inverts the safety ranking**: every arm errs in one direction
270
- only, Apple's errors are all `wait` where `stop` was right and both Qwen arms'
271
- are all `stop` where `wait` was right — and a wrong `stop` abandons a plan that
272
- would have worked, while a wrong `wait` costs a settle. A comparison reporting
273
- only the percentages would have recommended the wrong model.
274
-
275
- The Ollama arm is an **experiment, not a recommendation**: off unless named,
276
- no weights shipped, no dependency added, and every arm reads the same briefing
277
- out of `native/supervise.swift` so no arm is answering a different question.
278
-
279
- **Two further results, both negative, both worth more than the table.** A
280
- *cascade* — free rule first, model where it is unsure, Claude after that — was
281
- worse at every abstention band we tried (95% for the rule alone; 91%, 86%, 82%
282
- as more was handed to the model). And giving the model a fourth word, `abstain`,
283
- cost the Apple arm about a third of its accuracy (77/82/86% → 45/50/55%) **while
284
- it never used the word once**. One added paragraph, nothing else changed. So the
285
- fourth word ships off, and "add an abstain token" became "find a judge that will
286
- use one".
287
-
288
- The supervisor's whole vocabulary is three words on purpose. It cannot invent a
289
- step, skip one, substitute a target or continue past an unexpected screen — not
290
- because a threshold forbids it but because those are not answers it can give.
291
- That constraint replaced an earlier version of the same idea that was given
292
- latitude over *what* to open and pressed a button labelled "YES, THIS FIXED MY
293
- PROBLEM" in a live app.
294
-
295
- ## Capabilities are independent
296
-
297
- Each layer works without the ones above it, and `doctor` tells you which you
298
- have. **Observation needs nothing but Xcode.**
299
-
300
- | Capability | Needs | Without it |
301
- | --- | --- | --- |
302
- | Watch the screen, wait, recall | nothing extra | — |
303
- | Read labels + coordinates from pixels | `swiftc` (Xcode CLT) | falls back to the accessibility tree alone |
304
- | Tap, type, swipe | nothing extra | simframe observes but cannot touch |
305
- | Accessibility tree | nothing extra | OCR alone still yields labels and coordinates |
306
-
307
- `simframe doctor` names which engine is carrying each capability, per device.
308
- **Nothing beyond Xcode is required.** Capture, input, text recognition and the
309
- accessibility tree all run in-process, in one daemon.
310
-
311
- [`idb`](https://fbidb.io) is still accepted as a fallback for input and for the
312
- tree, for a machine where the daemon cannot run — and `SIMFRAME_AX_DRIVER=idb`
313
- forces the tree back onto it, which is the escape hatch if an Xcode upgrade
314
- breaks the host-side path.
315
-
316
- ```bash
317
- # optional fallback, not a requirement
318
- brew tap facebook/fb && brew install idb-companion && pipx install fb-idb
319
- ```
320
-
321
- ## Android
36
+ **Cursor** — [![Install in Cursor](https://cursor.com/deeplink/mcp-install-dark.svg)](https://cursor.com/en/install-mcp?name=simframe&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsInNpbWZyYW1lIiwibWNwIl19)
37
+ or add the same `simframe` block to `~/.cursor/mcp.json`.
322
38
 
323
- simframe's second backend drives an Android emulator with the same commands, the
324
- same screen map and the same memory as a simulator. Everything above the
325
- platform boundary — the frame store, settle, the structural fingerprint, the
326
- screen map, refs and the transition graph — runs on it unmodified, because the
327
- boundary hands it frames and nothing above it knows what a simulator is.
39
+ **Codex, Windsurf, Zed, VS Code, anything else that speaks MCP** — the same
40
+ JSON block, wherever that client keeps its `mcpServers`.
328
41
 
329
- | Capability | Android | How |
330
- | --- | --- | --- |
331
- | Watch the screen, wait, recall | yes | frames at **41 ms** through the emulator console, host-side — no adb in the capture path |
332
- | Read labels + coordinates from pixels | yes | the same Vision OCR + CV, off the same PNG |
333
- | Screen map, refs, screen memory, the graph | yes | unchanged above the boundary |
334
- | Tap, type, swipe, keys | yes | the console's `event mouse` as a real down/move/up, `event text` for characters, `input keyevent` for keys |
335
- | Clipboard, and `paste` into a field | yes | the emulator's gRPC `setClipboard`, over `node:http2`, no dependency, then `KEYCODE_PASTE` to deliver it |
336
- | List/resolve devices, launch, terminate, open a URL, permissions | yes | `adb`, with the permission state read back off the device |
337
- | Accessibility tree | **not available (OCR + CV only)** | `uiautomator dump` costs **2,012 ms** a read, against 45 ms for the iOS tree. See [`docs/DEFERRED.md`](docs/DEFERRED.md) |
338
- | A launch confirmed to have reached the front | **not available (the launch is not checked)** | iOS compares the pid `simctl launch` printed against the pid the device reports as frontmost, in **2–5 ms**. Nothing here reports either; `am start` fronts synchronously, which is why it has not bitten — but that is not a check. See [`docs/DEFERRED.md`](docs/DEFERRED.md) |
339
-
340
- ```bash
341
- # an emulator is found the same way a simulator is
342
- simframe devices # ● Small_Phone_API_36 Android 16 (API 36) emulator-5554
343
- simframe ui --device=emulator-5554
344
- simframe do --device=emulator-5554 flow.json
345
- ```
42
+ Then boot a simulator and ask the agent for something: *"open my app and tap
43
+ through the signup flow"*, *"is the list on the Orders tab loading?"*, *"why
44
+ does the Save button do nothing?"*. The first call builds a small Swift daemon
45
+ from source — about 15 seconds, once.
346
46
 
347
- A host with a booted simulator **and** a booted emulator has no default, and
348
- simframe will not pick one for you: preferring iOS because it came first would
349
- tap a simulator while you were driving an emulator, and acting on the wrong
350
- device is worse than refusing. So a command with no device names both and stops.
351
- `--device` answers it per command; `SIMFRAME_DEVICE` answers it per shell:
47
+ The CLI is the same thing without an agent in the loop, and the cheapest way
48
+ to check the install:
352
49
 
353
50
  ```bash
354
- export SIMFRAME_DEVICE=emulator-5554
355
- simframe ui # the emulator, without saying so every time
51
+ npm install -g simframe
52
+ simframe doctor # every layer, and which engine carries it
53
+ simframe ui # the current screen as a numbered text map
356
54
  ```
357
55
 
358
- The tree is a deliberate omission, not an oversight. Making it fast needs a
359
- resident instrumentation APK on the device — the shape uiautomator2, Maestro and
360
- Appium all converged on — and that would be simframe's first runtime artifact
361
- installed onto your device. The perception ladder was built so a missing tier
362
- degrades rather than fails, and this is exactly that case: OCR and CV yield
363
- labels and coordinates on Android today, and a tap by label works without a tree
364
- at all. `simframe doctor` reports the tier as `optional` with that number, so
365
- the gap is visible rather than silent, and the criteria for revisiting it are in
366
- `docs/DEFERRED.md` under **Phase 8b**.
367
-
368
- **What the missing tree costs, measured rather than hand-waved.** Screen
369
- *identity* is weaker on Android than on iOS, and specifically so. Tokens per
370
- screen, and where they come from:
371
-
372
- | Screen | Tokens | Regions | Roles | Chrome labels |
373
- | --- | --- | --- | --- | --- |
374
- | launcher | **1** | nav-bar 1 | text 1 | 0 |
375
- | Settings root | 9 | content 9 | text 9 | 0 |
376
- | example.com in Chrome | 6 | content 3, nav-bar 3 | text 6 | 0 |
56
+ If more than one simulator is booted, name the one you mean with
57
+ `--device=<udid>` or `export SIMFRAME_DEVICE=<udid>`; simframe refuses to guess,
58
+ because the command that would act on the wrong one is a tap.
377
59
 
378
- Every token has role `text`, because without a tree nothing infers a button from
379
- a rectangle reliably enough to say so, and no screen here carries a chrome label
380
- at all. So on Android a screen is recognised by the geometry of its text, which
381
- is thinner and noisier than the iOS mix of roles, chrome labels and geometry.
382
- Flows still work; screen *memory* is doing more guessing, and that is the honest
383
- cost of the tier being absent.
60
+ ## What it changes
384
61
 
385
- It is also why the obvious fix for the iOS drift — dropping content-region text
386
- out of identity, which would be a strict improvement there — is not available:
387
- it would leave Settings' root with zero tokens, and zero tokens is no identity
388
- at all. See [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
62
+ An agent driving a simulator is slow for three reasons, and only the first is
63
+ obvious. Every look is a wait (`simctl io screenshot` costs ~130 ms, paid on
64
+ every glance). Every step is a model round trip. And nothing is remembered, so
65
+ the same screen is re-read and re-reasoned about each time it appears. There
66
+ is a fourth that is pure waste: an image is the most expensive way to ask what
67
+ is on screen, and it still does not say what is tappable.
389
68
 
390
- That last column read **3** before this measurement changed it. Chrome's address
391
- bar is chrome by every structural test there is, so the screen's identity
392
- contained `"== example.com"` — a URL, meaning the same browser on a different
393
- page was a different screen and every learned route through it broke on
394
- navigation — plus `":"` and `"+"`, which are OCR reading punctuation off icons.
395
- A chrome label now has to be a name: two letters at minimum, and not an
396
- address.
397
-
398
- The emulator's own gRPC surface was checked for anything tree-shaped and has
399
- nothing: 43 RPCs for sensors, input, screenshots and VM state, and no notion of
400
- a view. That question is settled, not open. The same surface is what carries the
401
- clipboard.
402
-
403
- ## The tools
404
-
405
- Read first, act in batches, and look at pixels only when the question is about
406
- pixels. Every tool description says so, because a tool surface that does not
407
- steer the model is a tool surface the model uses wrong.
69
+ | | Without simframe | With simframe |
70
+ | --- | --- | --- |
71
+ | Look at the screen | ~130–400 ms screenshot, blocking | **~20 ms**, already captured |
72
+ | "Did anything change?" | a full image | **~2 ms**, text only |
73
+ | Reading a screen | an image, ~1,600 tokens, no tap points | **~330 tokens** of text, with tap points |
74
+ | A 10-step flow | 10 turns, 10 images | **1 turn, 0 images** |
75
+ | Same flow, again | every run is the first | replayed from memory, every step verified, **zero model calls** |
408
76
 
409
- | Tool | What it does |
410
- | --- | --- |
411
- | `sim_ui` | **Start here.** The screen as a numbered text map: region, type, label, state, tap point, source. A tenth the cost of a screenshot and strictly more useful. |
412
- | `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
413
- | `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
414
- | `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
415
- | `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
416
- | `sim_find` | Resolve an intent to one control, without acting on it. |
417
- | `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
418
- | `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
419
- | `sim_wait` | Waits for the screen to change *and then* settle. |
420
- | `sim_look` | **The only tool that returns an image**, capped at 1024 px. For layout, colour, spacing — questions text cannot answer. |
421
- | `sim_recall` · `sim_strip` | Look backwards: a text timeline of what happened, or recent frames tiled into one image. |
422
- | `sim_storage` | **What the app believes**, as opposed to what it drew: its `UserDefaults` and, for React Native, its `AsyncStorage`. Reads the data container off disk, so it answers on a device that is **not running**. |
423
- | `sim_capture` · `sim_devices` | Manage capture loops; list simulators. |
77
+ Measured on an iPhone 17 Pro simulator, iOS 26.5, Apple Silicon, against a real
78
+ production app; the full tables with N, median and p95 are in
79
+ [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
424
80
 
425
- ### What the screen looks like as text
81
+ This is what the agent gets back instead of a picture:
426
82
 
427
83
  ```
428
84
  iPhone 17 Pro · 402x874pt · screen a1b2c3d4 "Inbox" (known, 3 known exits)
@@ -437,745 +93,109 @@ content:
437
93
  tab-bar:
438
94
  #6 text 62,835 Inbox
439
95
  #7 text 201,835 Settings
96
+ next: settled; nothing ambiguous — chain the next steps in one sim_do without looking again.
440
97
  ```
441
98
 
442
- Region first, because "Inbox" the title and "Inbox" the tab differ only by where
443
- they are. A tap point, because that is what an action needs. And a number, which
444
- is a selector: whatever this calls `#3`, the next call can tap as `#3` without
445
- describing it. A ref is valid only while that screen is showing — used on a
446
- different screen it refuses rather than tapping whatever now sits there.
447
-
448
- `= something` is what the control *contains*, from the accessibility tree, and
449
- `~ something` is what OCR read off the pixels. Both are printed, and where they
450
- disagree that is the point: one is authoritative and the other is what is
451
- actually on screen, and a field mid-edit can legitimately differ. A row with no
452
- `=` is a control that reports no value, not an empty one.
453
-
454
- When the elements were recalled from screen memory rather than looked at just
455
- now, the header says so and how long ago — `elements recalled from 41s ago —
456
- pass refresh for what is there now`. Identity is cached on purpose, because a
457
- list with new rows is the same screen; contents are exactly what changes without
458
- the screen changing, so the age is worth seeing.
459
-
460
- Three ways to name a control, anywhere one is named:
461
-
462
- | | |
463
- | --- | --- |
464
- | `"Save"` · `the Assets tab` · `back` | **start here** — resolved by intent: verbs, typos, synonyms, icon-only controls by their common name |
465
- | `#3` | the number the map gave it. Cheap and exact, but only within the round trip that numbered it |
466
- | `@120,400` | raw point coordinates. Last resort: it cannot tell you it missed. |
99
+ Region, type, label, contents, a tap point, and a number the next call can use
100
+ as a selector. `= value` is what the accessibility tree says the control
101
+ contains; `~ value` is what OCR read off the pixels. Where they disagree, that
102
+ is the point.
467
103
 
468
- The order is deliberate and it used to be the other way round. Four peer rounds
469
- reported that intent resolution worked every time while refs renumbered
470
- underneath them, so a table that led with `#3` and called it "unambiguous" was
471
- recommending the more brittle of the two.
472
-
473
- ## Baselines: the thing to understand
474
-
475
- Every change question is really "changed **since when**?" — and the answer is
476
- almost never "since the previous frame". A UI transition is over in about 700 ms,
477
- so comparing consecutive frames tells a caller that polls every few seconds
478
- "nothing changed", even though the screen is completely different from when it
479
- last looked.
480
-
481
- So simframe compares against **the last frame you observed**. Over MCP that is
482
- automatic. From the CLI, capture a baseline before you act:
483
-
484
- ```bash
485
- H=$(simframe mark)
486
- # ...tap, launch, navigate...
487
- simframe wait --since=$H # change, then settle
488
- simframe state --since=$H # what moved, as text
489
- ```
490
-
491
- The same applies to waiting. `--mode=settle` (the default) waits for a change and
492
- *then* for stillness, because a bare "wait until stable" called in the moment
493
- before an animation starts will correctly, and uselessly, return immediately.
494
-
495
- ## Screen memory
496
-
497
- An accessibility tree is a promise apps do not always keep. In testing against a
498
- real production app, its custom tab bar published **no children at all**, its
499
- icon buttons carried unreadable private-use glyphs, and its React Native text
500
- inputs were **absent from the tree entirely** — the controls used most were
501
- exactly the ones that could not be tapped by name.
502
-
503
- So simframe reads the screen two ways and remembers the result:
504
-
505
- - **Accessibility** gives real hit targets, types and enabled state.
506
- - **On-device OCR** (Apple's Vision, ~290 ms, no model round trip) gives every
507
- label a person can actually see, with coordinates.
508
- - The merge is keyed by a **structural fingerprint**, so the next visit is a
509
- file read. What that fingerprint is, and why it is not a pixel hash, is below.
510
-
511
- ```
512
- first visit to a screen ~1000 ms read tree + OCR, store the map
513
- every visit after that ~1 ms look it up
514
- ```
515
-
516
- OCR is also more accurate than measuring by eye. On one tab bar the first tab
517
- centre sat at x=62, not the x=40 an even five-way split predicts — a silent
518
- mis-tap on every attempt.
519
-
520
- Two details that matter:
521
-
522
- - **Containers do not absorb their contents.** A tab bar encloses all five tab
523
- labels but is not any of them, so the merge only combines an element with text
524
- of comparable size.
525
- - **Ambiguity is reported, not guessed.** A word that is both a screen title and
526
- a tab returns an error listing both with coordinates, because silently tapping
527
- the title looks exactly like nothing happening.
528
-
529
- ### Two hashes, because there are two questions
530
-
531
- "Did this move?" and "is this the same screen?" look like one question and are
532
- not. simframe answers them separately, and getting that wrong was the single
533
- most expensive mistake in its development.
534
-
535
- **Change and settle** are questions about pixels, so a pixel hash answers them.
536
- The frame hash changes whenever any pixel group changes — a clock digit, one new
537
- row — which is exactly right for "did anything happen?" and useless as a key for
538
- "have I been here before?". For change detection there is a layout hash: status
539
- bar cropped, difference hash over a 12×24 grid.
540
-
541
- A mean-threshold hash was tried first and was actively dangerous: low-contrast
542
- screens collapsed onto identical values, so unrelated screens matched at distance
543
- 0 and taps landed on the wrong control. The difference hash fixed that.
544
-
545
- **Identity is not a question about pixels**, and this is the part that took three
546
- attempts. Content *is* pixels: a list whose rows changed drifts as far as a
547
- different screen does. Measured on a real app, same-screen revisits reached 62
548
- bits against a different-screen floor of 74 — overlapping, with no threshold
549
- available to choose. An earlier calibration had suggested a comfortable margin
550
- (0–4 against 77–113), but it was measured on screens whose content happened to be
551
- stable and did not survive contact with a real list.
552
-
553
- So identity is **structural**. The fingerprint is built from element roles,
554
- frames quantised to a 24 px grid, the region each element sits in, and repeated
555
- siblings bucketed as "one" or "many" rather than counted. Deliberately included:
556
- the labels of chrome elements only — nav title, tab labels, toolbar buttons —
557
- because two list screens with identical structure are told apart by their title
558
- and nothing else. Deliberately excluded: all content text and values, the status
559
- bar, and the keyboard region when a keyboard is up.
560
-
561
- It does not depend on the accessibility tree. Fingerprinting from OCR boxes
562
- alone, with the tree discarded entirely, still separates screens — different
563
- screens ceiling 0.35 against the same threshold.
104
+ ## The tools
564
105
 
565
- | | Jaccard similarity |
106
+ | Tool | What it does |
566
107
  | --- | --- |
567
- | Same screen, revisited | 0.41–1.00 |
568
- | **Different screens** | **0.00–0.31** |
569
-
570
- The threshold sits in that gap, but the gap is narrower than anyone would want,
571
- and one screen causes it: a screen whose sections load from different sources has
572
- more than one genuine settled structure, and two structures of one screen are as
573
- far apart as two different screens.
574
-
575
- No threshold can express that, so a screen may hold **several** accepted
576
- fingerprints instead. A new one is admitted only when a known edge lands
577
- somewhere its target does not recognise — the edge is the evidence that it is the
578
- same place — and only if no other stored screen claims that reading. Identity
579
- stays exact rather than being loosened, and the count is capped, so a
580
- non-deterministic action shows up as a node collecting variants rather than as
581
- screens silently merging.
582
-
583
- It earns its place on real apps: an app reconnecting to its bundler put an alert
584
- over one screen, and that screen gained a variant instead of a duplicate
585
- appearing.
108
+ | `sim_ui` | **Start here.** The screen as the numbered text map above. |
109
+ | `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. Fallback selectors, optional steps, and a bounded `seek` let a batch survive a surprise instead of handing it back. |
110
+ | `sim_state` | Has anything changed since your last look, and which regions moved. |
111
+ | `sim_goto` · `sim_flow_run` | Walk to a screen simframe has been to before; replay a saved flow with **zero model calls**. |
112
+ | `sim_find` · `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. |
113
+ | `sim_launch` · `sim_open_url` · `sim_permission` | Launch (confirmed to have reached the front), open a deep link, grant a privacy permission without tapping the alert. |
114
+ | `sim_look` | **The only tool that returns an image**, capped at 1024 px — for layout, colour and spacing. |
115
+ | `sim_recall` · `sim_strip` · `sim_storage` | What happened in the last minute, as text or as a contact sheet; what the app *saved* (`UserDefaults`, `AsyncStorage`), even on a device that is shut down. |
116
+ | `sim_wait` · `sim_capture` · `sim_devices` | Wait for change then settle; manage capture; list simulators and emulators. |
117
+
118
+ Three ways to name a control, anywhere one is named: `"Save"` (resolved by
119
+ intent — verbs, typos, synonyms, icon-only controls by their common name),
120
+ `#3` (the number the map gave it, valid for that screen only), or `@120,400`
121
+ (raw points, last resort). A label that matches the destructive vocabulary —
122
+ Delete, Pay, Send, Sign out — is never substituted for or guessed at.
123
+
124
+ A [Claude Code skill](skills/simframe/SKILL.md) ships in the package and
125
+ teaches the CLI path, which is cheaper still: a few lines of output instead of
126
+ a tool result, and no MCP schema in context until a tool is used.
586
127
 
587
- Failing to recognise a screen you have seen is harmless — it rebuilds the map and
588
- taps correctly. Matching the *wrong* screen taps the wrong control. The threshold
589
- is set to err toward the first.
590
-
591
- ## Navigating by memory
592
-
593
- Once simframe knows which screens exist and which action leads from one to the
594
- next, getting somewhere is a search over known edges rather than a question for a
595
- model:
596
-
597
- ```bash
598
- simframe screens # what this device has learned
599
- simframe goto invoices # walk there, verifying every step
600
- ```
601
-
602
- And when the screen and the behaviour disagree, the question is usually not
603
- about the screen at all:
604
-
605
- ```bash
606
- simframe storage # apps with a data container
607
- simframe storage com.example.myapp # what that app saved
608
- ```
609
-
610
- `sim_ui` says what is drawn; `sim_storage` says what the app believes. It reads
611
- the data container straight off the host filesystem, which means it works on a
612
- device that is **shut down** — `simctl` cannot do this at all, on any of its own
613
- paths, once a device stops running.
614
-
615
-
616
- Measured on a four-tab tour, `goto` plans and walks three-step routes with every
617
- step verified and no model call. It fails rather than guesses: an unknown
618
- destination, a query matching two screens equally, or no path of known edges all
619
- report themselves instead of tapping hopefully.
620
-
621
- Flows work the same way and refuse to save if any step went unverified —
622
- replaying a recording of something that may not have worked just reproduces the
623
- doubt.
624
-
625
- ```bash
626
- simframe flow save checkout ./checkout.json
627
- simframe flow run checkout
628
- ```
629
-
630
- ### What the memory is worth, isolated
631
-
632
- Three agent sessions drove the same task family on the same production app. The
633
- third existed only to hold a variable still: pass 2 beat pass 1 by twenty-one
634
- model round trips, but it had a warm graph *and* an operator who had already
635
- driven the app once, so the gap could belong to either.
636
-
637
- | pass | operator | graph | round trips |
638
- | --- | --- | --- | --- |
639
- | 1 | fresh | cold | **33** |
640
- | 3 | fresh | **warm** | **24** |
641
- | 2 | experienced | warm | **12** |
642
-
643
- **The graph is worth ~27%** of the round trips. Operator knowledge is worth the
644
- rest — half of everything left after the graph had taken its share. A second
645
- reporter, unable to isolate the graph, derived **~23%** from escalation rate per
646
- step without knowing that number.
647
-
648
- One app, one operator per pass, and round trips are not seconds. What it settles
649
- is the direction: the memory is real, it is the smaller half, and the larger half
650
- is what an operator learns and no tool here carries between sessions yet. Full
651
- working in [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) §14.
652
-
653
- ## Does this work on *your* app?
654
-
655
- Nothing in simframe is written for a particular app. What varies between apps is
656
- how much of the accessibility tree exists, and simframe is built to degrade
657
- rather than fail:
658
-
659
- - **Good tree** → tap by label, batch aggressively, everything just works.
660
- - **Partial tree** (custom tab bars, icon buttons) → OCR fills the gaps; you tap
661
- by the visible text instead.
662
- - **No tree at all** → OCR alone still yields labels and coordinates.
128
+ ## Android
663
129
 
664
- Run `simframe ui` on any screen to see exactly what simframe can see, with each
665
- target marked `ax` or `ocr`. If something you can read is not listed, that is a
666
- bug worth reporting.
130
+ The second backend drives an Android emulator with the same commands, the same
131
+ screen map and the same memory. Frames arrive at ~41 ms through the emulator
132
+ console; tap, swipe, text, keys and the clipboard work; `simframe devices`
133
+ lists emulators next to simulators. What Android does not have is an
134
+ accessibility tree, so its screens are read by OCR and CV alone — and
135
+ `simframe doctor` says so rather than pretending. Details, costs and the
136
+ reasoning are in [the guide](docs/GUIDE.md#android).
667
137
 
668
- A control the app **declares but never names** — an icon-only overflow menu, a
669
- back chevron — is listed with its coordinates and tappable by `#ref` rather than
670
- by name, and the map counts how many such controls are on the screen. A view the
671
- app never declared accessible at all is invisible to any accessibility tree, ours
672
- included, and the same line says so: a count of zero on a screen that has one
673
- would be the more expensive answer.
138
+ ## How it works, briefly
674
139
 
675
- **What a settle can and cannot see.** Stillness is decided from a mean over a
676
- grid of the frame, and a small animation does not move a mean. Measured on a
677
- screen built to never settle — a spinner, frames every 77–95 ms — the mean
678
- difference was **0.00196** against a 0.004 threshold while one cell moved by
679
- **0.0275**, and `stableForMs` reported **79 seconds** of stillness. So a settle
680
- can return satisfied while part of the screen is still moving, and it now says
681
- so: `settled after 63ms (a 13x13 region is still animating)`. A settle that
682
- times out names where the movement is instead. Refusing to settle on any
683
- animation would be right for a spinner and wrong for a blinking cursor, so the
684
- disagreement is reported rather than resolved by a guessed threshold.
140
+ One Swift daemon per device reads the simulator's framebuffer straight off its
141
+ IOSurface when the display reports damage, runs Apple's Vision OCR on the same
142
+ surface, reads the accessibility tree host-side, and sends taps as real
143
+ down/move/up sequences over the HID channel — no PNG, no file, no spawned
144
+ process. Frames land in `~/.simframe/<udid>/`, input goes over a `0600` Unix
145
+ socket, and the MCP server and CLI are thin clients over both.
685
146
 
686
- Two honest caveats. OCR reads **text**, so a purely graphical icon with no label
687
- is invisible to both paths — the tree still gives you its coordinates, and
688
- `#ref` still taps it. And the confirm-button vocabulary (`APPLY`, `OK`, `SAVE`,
689
- `DONE`…) is English; a localised UI needs those words extended.
147
+ Screen memory keys a merged element map by a structural fingerprint, so a
148
+ screen seen before is a file read. A transition graph records which action led
149
+ from which screen to which, so `sim_goto` plans a route and a saved flow
150
+ replays with no model in the loop.
690
151
 
691
- ## Measured
152
+ Every layer degrades rather than fails, and never quietly: `simframe doctor`
153
+ reports each one and `--strict` turns any downgrade into a non-zero exit. The
154
+ private frameworks the daemon links are documented, with evidence, in
155
+ [`docs/PRIVATE_API.md`](docs/PRIVATE_API.md).
692
156
 
693
- iPhone 17 Pro, iOS 26.5, Apple Silicon, default settings.
157
+ ## Read on
694
158
 
695
159
  | | |
696
160
  | --- | --- |
697
- | Frame capture, whole pipeline | **6.6 ms** |
698
- | Frame grab alone | **0.13 ms** |
699
- | The `simctl` + `sips` path it replaces | ~210 ms |
700
- | Warm frame read (`sim_look`) | ~20 ms |
701
- | State check (`sim_state`) | ~2 ms |
702
- | Input round trip (`ping`) | **0 ms** |
703
- | Tap (70 ms hold / 10 ms hold) | 76 ms / 13 ms |
704
- | Text recognition, in-process | **~174 ms** |
705
- | Text recognition, via PNG + helper (fallback) | ~555 ms |
706
- | Accessibility tree read, in-process | **~45 ms** |
707
- | Accessibility tree read, via idb (fallback) | ~203 ms |
708
- | Screen map: first visit / remembered | ~305 ms / **~1 ms** |
709
- | CPU | 1.1 % idle · 3.1 % active |
710
- | Frame memory | ~60 s of screen, ~2.7 MB |
711
-
712
- Reproduce all of it with `npm run bench`, which prints the same table against
713
- your machine. Full detail, including the measurement traps, is in
714
- [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
715
-
716
- ### Wall clock per step — where the time really goes
717
-
718
- Every number above is microscopic next to the one that decides how fast this
719
- feels, and it took two field reports to see it. Per-step wall clock is
720
-
721
- ```
722
- (model round trip + simframe work × n) / n for n steps in one call
723
- ```
724
-
725
- **Compare like with like, which this table used not to.** A human tester's
726
- 1.95 s per step is the *whole* loop — look at the screen, decide what to do,
727
- do it. Any row that does not include a decision is not comparable to it.
728
-
729
- | | perceives | decides | acts | per step |
730
- | --- | --- | --- | --- | --- |
731
- | **a human tester, measured** | yes | yes | yes | **1.95 s** |
732
- | **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
733
- | simframe, batch of 4 | yes | yes | yes | ~6.7 s |
734
- | simframe, one model call per step | yes | yes | yes | ~21.7 s |
735
- | — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
736
- | — a saved flow replayed | yes | **no** | yes | 1.98 s |
737
-
738
- **The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
739
- a human**, and about eleven times when a failure forces one model call per step.
740
-
741
- The two indented rows are the ones this README used to lead with, and both are
742
- category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
743
- taken out — the thinking is the model round trip, which is most of the clock.
744
- And a replay *decides nothing*: it is a recording being played back, so its fair
745
- counterpart is a human repeating a flow they have memorised, who would be well
746
- under 1.95 s. Replay against a human working something out for the first time is
747
- a rehearsal measured against a first attempt.
748
-
749
- What the two rows do establish, and it is the finding that reordered this
750
- project: **there is nothing left to win inside the engine.** 1.7 s is small
751
- beside a 20 s round trip, so making perception or input faster buys single-digit
752
- percentages. The only variable that matters is `n` — how many steps one decision
753
- covers. Every improvement here has come from raising it, not from faster code.
754
-
755
- ### One more term: the launch is a fixed cost
756
-
757
- Measured on the benchmark suite with **no model in the loop at all**:
758
-
759
- | route | steps | agent | per step | human | per step |
760
- | --- | --- | --- | --- | --- | --- |
761
- | contacts-kate-bell | 2 | 12419 ms | **6.2 s** | 4300 ms | 2.15 s |
762
- | settings-larger-text | 4 | 15281 ms | 3.8 s | 7799 ms | 1.95 s |
763
-
764
- 2.9× a human on the short route with nothing thinking, which does not fit
765
- `1.7s × n`. A cold app launch is a **one-off cost amortised over the route**:
766
-
767
- ```
768
- per step = (model round trip + launch cost + ~1.7s × n) / n
769
- ```
770
-
771
- That reconciles a 7-step replay at 1.82 s/step with a 2-step route at 6.2 s/step
772
- — one launch spread over 7 steps or over 2. **The `~1.7 s` figure is warm taps
773
- inside a batch**, and short routes are materially worse than the table above
774
- implies on its own.
775
-
776
- `step_ratio` was **1** throughout: when a run completed it took exactly the
777
- minimum number of steps.
778
- A field report put the split at **34% simframe, 60% agent round trips** over
779
- 462 s of wall clock — the tester's *"30+ seconds between each step"* was
780
- accurate and was not simframe. So there is nothing left to win inside the
781
- engine, and the only variable is `n`: `simframe hpi` reports `steps_per_call`
782
- for exactly that reason.
783
-
784
- Which is why **every hard-fail that drops a caller back to single-stepping is a
785
- latency bug**. Between two field reports on the same flow, one run took 33 tool
786
- calls and the next took **16**, for 28 executed steps and no screenshots at
787
- all — the difference being defects fixed, not code made faster.
788
-
789
- ## How it works
790
-
791
- ```
792
- ┌─── simframed — one Swift daemon per simulator ──────────────────────────────┐
793
- │ │
794
- │ display damage callback ──► read IOSurface ──► scale ──► hash │
795
- │ the screen tells us 0.13 ms 6.6 ms total │
796
- │ │ │
797
- │ Vision OCR reads the same surface ─────┤ no PNG, no file, no spawn │
798
- │ │ │
799
- │ Indigo HID ◄── control socket ◄────────┤ 0600, one JSON object per line │
800
- │ taps, swipes, text │ │
801
- │ ▼ │
802
- │ ~/.simframe/<udid>/ frames · state.json · meta.json │
803
- └─────────────────────────────────────────────────────────────────────────────┘
804
- │ a rename is atomic
805
- ┌────────────────────────────────────▼────────────────────────────────────────┐
806
- │ MCP server / CLI: stat + read, or one socket round trip for input. │
807
- │ Screen memory: layout hash ──► label → point, built once per screen. │
808
- └─────────────────────────────────────────────────────────────────────────────┘
809
- ```
810
-
811
- The daemon links CoreSimulator and SimulatorKit, which are private frameworks
812
- with no documentation and no stability promise. Everything it calls is recorded
813
- in [`docs/PRIVATE_API.md`](docs/PRIVATE_API.md) with the evidence behind it, so
814
- an Xcode upgrade that moves something is a bounded fix rather than an
815
- archaeology project. If a layer breaks, simframe degrades to the layer below
816
- and `doctor` says which.
817
-
818
- Run `simframe start --engine=screenshot` to use the one-frame-at-a-time loop
819
- instead — it is also the only capture engine on Android, where it reaches the
820
- emulator console rather than any simulator tool. `--engine=simctl` is still
821
- accepted as the name that loop used to have.
822
-
823
- - **Files are the IPC for reads.** The daemon renames completed frames into
824
- place; readers just read them. A rename is atomic, so a reader can never see a
825
- half-written frame. Input is the one thing that needs a reply, and it goes
826
- over a `0600` Unix socket — the file system is the whole permission model.
827
- - **Almost no dependencies.** The only runtime npm dependency is the MCP SDK.
828
- The daemon is Swift built from source against frameworks already on the
829
- machine.
830
- - **The screen says when it changed.** The capture loop is driven by the
831
- display's damage callback rather than a timer, so an idle screen costs
832
- nothing and a moving one is picked up at once.
833
- - **It backs off when nothing happens.** 4 fps while the screen moves, 1.5 fps
834
- once still, snapping back instantly on change.
835
- - **One writer per device.** Ownership lives in `meta.json`; `stop` refuses to
836
- kill a loop another client is using unless forced.
837
- - **A wedged capture loop never looks like a calm screen.** Every answer carries
838
- a liveness check, and `wait` fails loudly rather than quietly timing out.
839
- - **An action with no visible effect is reported, not waited out.** Selecting a
840
- radio button moves ~0.1 % of the screen — below the change threshold — which
841
- used to burn the full timeout. Now the step returns in ~3 s marked
842
- `[no visible change]`, so you know to check rather than wait.
843
- - **And a gesture aimed at a coordinate says what it landed on.** A tap or swipe
844
- that changed nothing prints the element covering its start point, because a
845
- gesture goes to whatever is on top there:
846
-
847
- ```
848
- swiped 83,141 -> 83,800 [no visible change]
849
- [the swipe start point 83,141 is inside "Settings" (nav-bar)]
850
- ```
851
-
852
- Reported from the field as fifteen minutes lost to six identical swipes that
853
- a support banner was swallowing — a banner that was *in the element list the
854
- same call printed*. Nothing said "there is something at y≈753 and you started
855
- at y=750".
856
- - **A control half off the edge is offered, clamped, and says so.** A filter
857
- chip with 29 pt of itself visible used to be dropped — the off-screen test
858
- reads the *centre* — and OCR's reading of that same sliver, a box labelled
859
- `Flc`, was printed in its place. So the map did not merely omit a control; it
860
- offered a meaningless name for it at an ordinary-looking coordinate. Now:
861
-
862
- ```
863
- #6 button 387,191 Flowering (partly off-screen — the coordinate is the middle of the visible part)
864
- ```
865
-
866
- `scrollTo` is deliberately unchanged: a clipped element is still "not in
867
- view", so it still scrolls rather than calling a sliver good enough.
868
-
869
- ## CLI
870
-
871
- The CLI is the low-token path, and it is a first-class one: `--json` is on every
872
- command, so nothing has to be parsed out of prose.
873
-
874
- ```bash
875
- simframe ui # the numbered screen map — start here
876
- simframe ui --json | jq '.elements[] | select(.type=="button") | .label'
877
- simframe do flow.json # a whole flow, verified, then the end-state map
878
- simframe do flow.json --save=checkout # save it if every step verified
879
- simframe flow run checkout # replay it
880
- simframe tap "#3" # or "Save", or "@120,400"
881
- simframe find "the save button" # resolve an intent without acting
882
- simframe screens # screens this device has learned
883
- simframe goto invoices # walk to a known screen over known steps
884
- simframe state --since=$H # what changed, as text
885
- simframe mark # hash of the current frame, for --since
886
- simframe wait --since=$H # change, then settle
887
- simframe recall # what happened in the last minute, as text
888
- simframe recall --ago=15000 # the frame from 15s ago
889
- simframe frame --out=now.png # newest frame, native resolution, to a file
890
- simframe strip --count=6 # contact sheet, for an animation
891
- simframe doctor --strict # any degraded layer is a non-zero exit
892
- simframe escalations # why simframe still needs a model, by reason
893
- simframe escalations --session # ...this agent only, not every agent on the device
894
- simframe supervisions # local supervisor rulings, and what came of each
895
- simframe hpi # speed and accuracy against a human baseline
896
- simframe baseline record settings-larger-text --runs=5 # record the human
897
- simframe input reset # rebuild the HID session, without restarting anything
898
- simframe diagnose # what this device is doing right now, and which failure it is
899
- simframe revive # power-cycle a wedged device: stop, shutdown, boot, start, reset input
900
- simframe start / status / stop [--force] / devices
901
- simframe ui --device=emulator-5554 # or export SIMFRAME_DEVICE once
902
- ```
903
-
904
- `--device` takes `--device=X` and `--device X` alike. It used to take only the
905
- first: the space form set the flag to `true` and then resolved a device named
906
- "true", which is a poor answer to a flag `doctor`'s own advice tells you to
907
- type.
908
-
909
- ### When the simulator stops rendering
910
-
911
- A simulator driven hard for several minutes can stop rendering: capture fails,
912
- `simctl screenshot` fails too, and the daemon's own recoveries — re-resolving the
913
- display port, then rebinding the device — do not help. It reports the state
914
- rather than acting on it, because a capture loop that rebooted the device it was
915
- watching would be a tool reaching for the mains:
916
-
917
- ```
918
- capture is wedged and both recoveries are spent (2 port re-resolves, 2 device
919
- rebinds, no frame since). This needs the device restarted —
920
- `simframe revive --device=<udid>`. Backing off until a frame arrives.
921
- ```
922
-
923
- **`simframe diagnose` says which failure it is**, which `doctor` cannot: `doctor`
924
- answers "can this machine capture", and this answers "what is this device doing
925
- right now". It reads only what discriminates — frame sequence, age and
926
- stillness, element counts split by sensor, how many of them fuse, and who holds
927
- the front by pid — and returns one of:
928
-
929
- | verdict | what it means |
930
- | --- | --- |
931
- | `capture-down` | no frames at all, carrying the daemon's own sentence |
932
- | `nothing-readable` | frames arriving, neither sensor finds a single element |
933
- | `stale-frame` | both sensors full, almost nothing fuses — the framebuffer is behind the tree, so **an image from this device is not safe to trust** |
934
- | `not-presenting` | an app holds the front by pid and the display shows almost nothing |
935
- | `healthy` | — |
936
-
937
- `not-presenting` deliberately does **not** say whether that is a lock screen, a
938
- dead surface or a crashed system shell. It is not knowable from here, and
939
- guessing is how a regex ended up standing where a measurement belongs.
940
-
941
- The `stale-frame` threshold is measured rather than chosen: element fusion on
942
- five healthy screens ran 0.667–0.929, so the threshold sits at 0.1 — 6.7× below
943
- the observed floor rather than inside the metric's own noise. Numbers in
944
- [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
945
-
946
- `simframe revive` is that restart, in the order that matters — stop the daemon,
947
- shut the device down, boot it and *wait for the boot to finish*, start capture,
948
- rebuild the HID session — and it ends by checking frames are flowing again
949
- rather than by reporting that the steps ran. It is a command and not a
950
- behaviour: the decision stays yours.
951
-
952
- ### Reading what the local supervisor decided
953
-
954
- With `SIMFRAME_SUPERVISOR=apple`, every consultation is written to
955
- `~/.simframe/<udid>/supervisions.jsonl` with what the executor observed
956
- afterwards, which is what makes a ruling scoreable rather than merely recorded.
957
- `simframe supervisions` reads it:
958
-
959
- ```
960
- 14 supervisor rulings on iPhone 17 Pro
961
-
962
- stop -> stopped 7
963
- wait -> recovered 5
964
- wait -> still_failed 2
965
-
966
- sourced: model 14
967
- median latency: 1440ms
968
- edges the graph had timed: 0/14 — 14 ruling(s) are on edges with no p95
969
- ```
970
-
971
- That last line is the honest one: a step that failed is usually a step that has
972
- never succeeded on that edge, so the graph has no timing to compare against.
973
-
974
- ### Keeping a session cheap
975
-
976
- The expensive part of driving a simulator with an agent is not the tapping, it
977
- is the thinking between taps — observe, think, tap, observe, think. Measured
978
- over one real session against a third-party app: **62 tool calls for 179 steps**,
979
- and 48 of those calls were three steps or fewer. A twelve-step flow arrived as
980
- five calls, and every boundary between them was a think.
981
-
982
- Three things move that number, and simframe does the first two for you:
983
-
984
- - **Batch.** `sim_do` runs a whole flow in one call, with an assert after each
985
- step that matters. The asserts are what make it safe not to look in between:
986
- a step that lands somewhere unplanned halts the flow instead of letting the
987
- next four run against the wrong screen.
988
- - **A `next:` line on every action result** — from the CLI and the MCP server
989
- alike — computed locally from what the daemon already knows — whether the screen settled, whether the graph
990
- recognises it, how many elements it has, whether any labels repeat. When it
991
- says *nothing ambiguous — chain the next steps in one sim_do without looking
992
- again*, that is the tool telling the agent it does not need to think.
993
- - **A trailing map that was re-read, not recalled.** An action pays one
994
- perception pass — a few hundred milliseconds, locally — so the map it returns
995
- is the screen as it is now. The alternative was an agent spending a whole turn
996
- on `ui --refresh` because it could not trust the one it was given.
997
- - **One goal per session.** Sessions get slower with every turn. A flow that
998
- runs as one call adds one exchange to the context instead of twelve.
999
-
1000
- And one thing to know about waiting: `settle` asks whether the screen stopped
1001
- moving, and a screen waiting on a network call has stopped moving. For anything
1002
- that arrives over the network, assert on the content you expect —
1003
- `{"waitFor": {"value": "Kate Bell"}}` — rather than on stillness. The map says
1004
- `STILL LOADING` when the classifier can see a load in flight, but only you know
1005
- what "arrived" means.
1006
-
1007
- And the expensive habit worth naming: in that session, **28 of 62 calls returned
1008
- a screenshot** — about a third of its entire token cost — because the text map
1009
- could not report what a text field contained. It can now, so check the map
1010
- before reaching for pixels: a row carries the element's contents (`= Fryer 3`)
1011
- and its state (`disabled`), and the flow's own verdict already said whether the
1012
- action worked.
1013
-
1014
- ### The Claude Code skill
1015
-
1016
- [`skills/simframe/SKILL.md`](skills/simframe/SKILL.md) teaches the CLI path
1017
- directly: the cheap-to-expensive order, the selector grammar, what each verdict
1018
- means and what to do about it. It ships with the package, so an installed copy
1019
- has it.
1020
-
1021
- ```bash
1022
- mkdir -p ~/.claude/skills
1023
- ln -s "$(npm root -g)/simframe/skills/simframe" ~/.claude/skills/simframe
1024
- ```
1025
-
1026
- A skill and an MCP server are not redundant. The MCP server is discoverable —
1027
- it appears in the tool list without anybody setting it up. The skill is cheaper:
1028
- Claude reads 3–5 lines of CLI output instead of a tool result, and none of the
1029
- MCP schema is in context until a tool is actually used. Ship both, use whichever
1030
- the client makes easy.
1031
-
1032
- ## Degrading is allowed. Degrading quietly is not
1033
-
1034
- simframe is built to degrade rather than fail: no Swift toolchain still gives
1035
- you frames through `simctl`, no accessibility tree still gives you OCR. That
1036
- policy is right, and it nearly sank the tool twice — because a downgrade looked
1037
- exactly like everything working.
1038
-
1039
- Once, one file was missing from the published package, so `Package.swift`
1040
- declared a test target with no directory, SwiftPM reported overlapping sources,
1041
- and **every install silently fell back to the slow engine**. Another time OCR
1042
- shipped disabled the same way. Both passed the tests. Both printed nothing. The
1043
- bug was never the missing file; it was the silence.
1044
-
1045
- So every downgrade now announces itself:
1046
-
1047
- - `simframe start` prints the engine it chose, and if it is the slow one, why —
1048
- build error, missing sources, or "reason unrecorded" if the daemon was started
1049
- by an earlier process.
1050
- - `simframe doctor` marks a degraded layer `WARN`, not `ok`, and summarises what
1051
- is degraded and what that costs.
1052
- - A dependency that is simply not installed is `--`, not `WARN`. The distinction
1053
- is deliberate: `WARN` means this machine could be doing better and silently is
1054
- not, which is the failure worth shouting about. An optional fallback missing on
1055
- a fresh machine has not degraded from anything, and `--strict` ignores it.
1056
- - A device whose capture has **wedged** says so: `capture: stalled — the display
1057
- surface has been unreadable for 62s; 3 re-attaches did not help; only
1058
- restarting the device is known to cure it`. This is a different thing from a
1059
- still screen, and it used to look identical, because a damage-driven engine
1060
- produces no frames for either. An agent told "nothing changed" keeps tapping;
1061
- one told the simulator is wedged stops. simframe reports it and does not
1062
- restart your device.
1063
- - `--strict`, or `SIMFRAME_STRICT=1`, turns any downgrade into a non-zero exit.
1064
- CI runs strict, so a release cannot ship in the state that shipped twice.
1065
-
1066
- ```
1067
- $ simframe doctor
1068
- ok capture engine simframed
1069
- WARN input driver idb — the daemon's control socket is not up
1070
- WARN accessibility tree idb — the host-side translator did not load
1071
- ```
1072
-
1073
- Two checks enforce it. A packaging check derives the required file list from the
1074
- build's own inputs — a hand-written list is what rotted last time — and runs in
1075
- seconds without a simulator. An integration job installs the packed tarball on a
1076
- real simulator and asserts `capture.engine`, `input.driver` and `ocr.available`
1077
- are all the good values, under `--strict`.
161
+ | [`docs/GUIDE.md`](docs/GUIDE.md) | the full guide: recovering without a round trip, screen memory and the two hashes, navigating by memory, the CLI, diagnosing a wedged simulator, keeping a session cheap, the roadmap |
162
+ | [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | every number, with N, median, p95, machine and the measurement traps |
163
+ | [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
164
+ | [`docs/DECISIONS.md`](docs/DECISIONS.md) | the judgements that changed the plan, including a phase cancelled by its own measurement |
165
+ | [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
166
+ | [`docs/ARTICLE.md`](docs/ARTICLE.md) | *Agents shouldn't blink* — the argument the whole thing adds up to |
1078
167
 
1079
- That last check found a real bug the day it was written: a daemon shutting down
1080
- unlinked the control socket unconditionally, so restarting deleted the *new*
1081
- daemon's socket. Capture kept working, input quietly dropped to idb, and nothing
1082
- said a word — the exact failure shape, found by the thing built to catch it.
168
+ Almost everything here was decided by a measurement rather than an argument,
169
+ and several measurements reversed a decision that had already shipped. Two
170
+ headline results: a warm transition graph removes about 27% of an agent's
171
+ model round trips, and a one-line threshold on a number the daemon already
172
+ computes beat every local model tried as a step supervisor — 95% against
173
+ 77–91%, at zero latency.
1083
174
 
1084
175
  ## Limitations
1085
176
 
1086
- - Simulators and Android emulators only. Neither the framebuffer nor `simctl`
1087
- nor the emulator console can reach a physical device.
1088
- - Android has no accessibility tree, so its screen identity rests on the
1089
- geometry of OCR'd text: thinner and noisier than iOS's. See
1090
- [Android](#android) above.
1091
- - The daemon depends on private frameworks. They are stable enough to build on —
1092
- capture and accessibility survived the iOS 26 transition — but an Xcode
1093
- upgrade can move a symbol. `doctor` reports each layer separately so a break
1094
- is visible rather than mysterious, and `--engine=screenshot` still works.
1095
- - Hardware buttons: only `home` is implemented. The other Indigo codes are
1096
- unverified, and a wrong one can crash `backboardd` or lock the device, so they
1097
- return an error rather than a guess.
1098
- - Typing sends key positions, which iOS maps through the device's active
1099
- keyboard layout. Text that must be exact goes through the pasteboard, which
1100
- `sim_do` does by default.
1101
- - Region maps need a baseline inside the ~90 s history window. Older baselines
1102
- still get a reliable changed / did-not-change, without a map of what moved.
1103
- - Screen memory assumes a screen's layout is stable. A screen that reflows
1104
- dramatically between visits will simply be rebuilt.
1105
- - It speeds up *confirming* a fix, not *locating* one. A bug living in a memo
1106
- comparator or a stale closure is not visible in any frame.
1107
- - A switch is tapped at its **activation point** when the app publishes one —
1108
- UIKit's `accessibilityActivationPoint`, which is what a switch answers with,
1109
- and tapping it flipped a real switch 3 of 3 times where the frame centre
1110
- managed 0 of 3. Only **4 of 75** elements on a measured screen publish one,
1111
- though, so where the app says nothing the tap still goes to the centre of the
1112
- frame — and a switch's frame is the whole row, so it lands on the label and
1113
- the control at the trailing end does not move. `@x,y` remains the escape
1114
- hatch there. Never a guessed offset: nil means the app did not answer.
1115
- - A launch is confirmed to have fronted **on iOS only**. It compares the pid
1116
- `simctl launch` printed against the pid the device reports as frontmost, in
1117
- 2–5 ms. Android reports neither, so a launch there is not checked — `am start`
1118
- fronts synchronously, which is why it has not bitten, but that is not a check
1119
- and `doctor` says so. Note that a launch which starts a process without
1120
- bringing it forward now **fails** rather than returning success: a flow that
1121
- used to pass through such a launch and then assert on the previous app's
1122
- screen will start failing, correctly.
1123
- - The simulator's display pipeline stops rendering under rapid app relaunch —
1124
- about six cycles, reproducibly — and every frame comes back black while
1125
- `simctl` itself reports success. simframe now says so instead of reading a
1126
- black screen as a calm one, but it cannot fix it: restarting the device is
1127
- the cure that always works, and it usually recovers on its own.
1128
-
1129
- ## What was decided, and what was not built
1130
-
1131
- [`docs/DECISIONS.md`](docs/DECISIONS.md) is the register of judgements that
1132
- changed the plan: a phase cancelled by its own measurement, two features built
1133
- and reverted for cause, and the premises that turned out to be false. It is
1134
- short on purpose — the numbers live in
1135
- [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) and the open work in
1136
- [`docs/DEFERRED.md`](docs/DEFERRED.md).
1137
-
1138
- The most useful entry is a **no-go**: Phase 17 proposed a small on-device model
1139
- to choose the next element, and measuring the prize before the model showed the
1140
- matcher already resolves 37 of 40 real decisions. By the time a step reaches
1141
- simframe the decision has already been made — the goal names the option,
1142
- because the agent chose it and then asked for it by name.
1143
-
1144
- ## Roadmap
1145
-
1146
- - **Extend the confirm vocabulary beyond English.**
1147
- - **Region bands from clustering**, replacing the positional bands. They have
1148
- produced three bugs in three phases, and on Android they put a URL bar in the
1149
- nav bar and its URL into the screen's identity.
1150
- - **Phase 8b, conditionally:** an instrumentation APK for the Android
1151
- accessibility tree, with the criteria for doing it stated in
1152
- `docs/DEFERRED.md` rather than left to enthusiasm.
177
+ - Simulators and emulators only; nothing here can reach a physical device.
178
+ - The daemon links private frameworks. They survived the iOS 26 transition,
179
+ but an Xcode upgrade can move a symbol; `doctor` names the broken layer and
180
+ `--engine=screenshot` still works.
181
+ - Android has no accessibility tree, so screen identity there rests on OCR
182
+ geometry alone: thinner and noisier than iOS.
183
+ - Hardware buttons: only `home`. The other codes are unverified and a wrong
184
+ one can crash the device, so they return an error rather than a guess.
185
+ - The confirm vocabulary (`OK`, `Save`, `Done`…) is English.
186
+ - It speeds up *confirming* a fix, not *locating* one. A bug in a stale
187
+ closure is not visible in any frame.
188
+
189
+ The full list, with the measurements behind each item, is in
190
+ [the guide](docs/GUIDE.md#limitations).
1153
191
 
1154
192
  ## Releasing
1155
193
 
1156
- `npm version` runs a `version` hook that rewrites `server.json` to match and
1157
- stages it, so one command covers both files:
1158
-
1159
- ```bash
1160
- npm version minor # bumps package.json + server.json, commits, tags
1161
- git push --follow-tags
1162
- ```
1163
-
1164
- Before that hook existed, `server.json` had to be hand-edited between two
1165
- commands, and the release that forgot failed at the workflow's own agreement
1166
- check — which is the one thing that check is for.
1167
-
1168
- The `release` workflow verifies tag/`package.json`/`server.json` agree, validates
1169
- `server.json` against the live registry, and publishes to npm and the MCP
1170
- Registry. It holds **no secrets** — both halves authenticate with the workflow's
1171
- GitHub OIDC identity.
1172
-
1173
- That needs one setup step on npmjs.com, not in this repo: the package must have a
1174
- Trusted Publisher pointing at this repository and `release.yml` (Package →
1175
- Settings → Trusted Publisher → GitHub Actions). Without it npm has nothing to
1176
- trust and fails with `ENEEDAUTH`. npm is ending token publishing in January 2027,
1177
- and the tokens that work in CI need 2FA bypass, which npm's own UI warns against —
1178
- so OIDC is the durable path, not merely the tidier one.
194
+ `npm version` is the only way to bump a version — its hook keeps `server.json`
195
+ in step — and `git push --follow-tags` runs a workflow that publishes to npm
196
+ and the official MCP Registry with no secrets. Why it is built that way, and
197
+ the release that once vanished for a week, is in
198
+ [the guide](docs/GUIDE.md#releasing).
1179
199
 
1180
200
  ## License
1181
201