simframe 0.10.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -18,6 +18,193 @@ and CV alone. Tapping by label works; screen recognition is thinner, so prefer
18
18
  naming a device explicitly and re-reading the screen after a step you are unsure
19
19
  about.
20
20
 
21
+ ## The protocol: plan once, execute once, think only when told to
22
+
23
+ The expensive thing in a simulator session is not the tapping. It is you —
24
+ observe, think, tap, observe, think. Measured over a real session: 62 tool
25
+ calls for 179 steps, and 48 of those calls were three steps or fewer. A
26
+ twelve-step flow arrived as five calls, and every boundary between them was a
27
+ think that nothing had asked for.
28
+
29
+ So the protocol is four steps, and step 4 is the one that saves the time.
30
+
31
+ 1. **State the goal, then read the screen once.** `simframe ui`, or `sim_find`
32
+ if you doubt a selector will resolve. Once — not per step.
33
+ 2. **Write the entire flow as one `sim_do`**, with an assert after each step
34
+ whose success you would otherwise have checked by looking.
35
+ 3. **Run it. Read only the verdict line per step.**
36
+ 4. **Think again only when the result tells you to.** Every action result ends
37
+ with a `next:` line the daemon computed locally, and it says which case you
38
+ are in:
39
+
40
+ - `next: settled; screen known (…); N elements; nothing ambiguous — chain the
41
+ next steps in one sim_do without looking again` → **do not look. Act.**
42
+ - `next: new screen, nothing predicted here yet` → read it before acting on a
43
+ label you have not seen on it.
44
+ - `next: N labels repeat on this screen` → address those by `#ref`.
45
+ - `next: the flow stopped here` → this is the moment to think.
46
+ - `worked here before: tap "…" (7x)` → the graph's own vocabulary for this
47
+ screen, most-used first. It is evidence for writing a chain, not an
48
+ instruction: the most-tapped control on a screen is often what earlier runs
49
+ used to *back out*.
50
+ - `memory disagrees with this screen: N remembered controls not present` →
51
+ two screens share one fingerprint. Trust the element list, not the graph,
52
+ and re-read before anything irreversible.
53
+
54
+ **Do not** narrate each step, re-read the screen after every action, or use
55
+ extended thinking inside a flow. The flow is already planned; executing it is
56
+ not a decision.
57
+
58
+ ### Recovering without a round trip
59
+
60
+ Three things let a batch survive a problem instead of handing it back. Every one
61
+ of them exists because a real run lost a call to the thing it prevents.
62
+
63
+ **Fallback selectors.** `{"tap": "Save", "or": ["Done", "Confirm"]}` — tried
64
+ locally, in order, and only an exhausted list reaches you. Eligible after a
65
+ selector that did not *resolve*, and nothing else: retrying from a screen you did
66
+ not expect to be on is a second guess, not a retry. simframe refuses to
67
+ substitute a label that looks destructive even when you list it.
68
+
69
+ They are a cure for *"I named it wrong"*, and most real failures are *"it is not
70
+ there yet"* — so when everything in an `or` chain misses and the screen has only
71
+ just stopped moving, the failure says so. Reach for `waitFor` there, not for more
72
+ labels.
73
+
74
+ **`{"seek": "change username", "budget": 6}`** looks for something that is not on
75
+ this screen: it opens containers, checks, and comes back, depth first.
76
+
77
+ It **acts** — opening a door changes state — and it refuses to open anything that
78
+ commits, abandons or answers. It once opened `CANCEL` and pressed *"YES, THIS
79
+ FIXED MY PROBLEM"*, which is why that sentence is here. It does not tap the
80
+ target: it leaves you on the screen where the target resolves, and returns to
81
+ where it started if it fails. **Do not point it into a flow whose progress you
82
+ cannot afford to lose.**
83
+
84
+ **`sweep` — read and fill a long screen.** A viewport is the only honest unit:
85
+ the tree publishes what is rendered, so a form taller than the screen is knowable
86
+ only in pieces.
87
+
88
+ ```json
89
+ {"sweep": "all", "fill": {"Last Name": "Asadi", "Email": "a@b.c", "Comment": "…"}}
90
+ ```
91
+
92
+ It goes to the top, then reads section by section to the bottom, filling each
93
+ field **while it is on screen** — which beats finding a field and then trying to
94
+ scroll back to it, because one gesture travels a non-deterministic distance.
95
+ `{"from": "here"}` sweeps down from where you are. `{"sweep": "<text>"}` stops as
96
+ soon as it finds that text. It reports which section each element was in, what it
97
+ filled, and what it never found at any scroll position.
98
+
99
+ Prefer it to `scrollTo` on any form or long list. `scrollTo` hunts one label and
100
+ cannot work when the label is not yet rendered; a sweep covers the screen.
101
+
102
+ **The local supervisor**, when one is enabled, decides whether a failed step
103
+ should `wait`, `retry` or `stop` — before the failure reaches you. It knows
104
+ nothing about the app and you do, so brief it from the plan:
105
+
106
+ ```json
107
+ {"supervisor": "apple",
108
+ "supervise": "Lists here render a count header before their rows, so a missing row usually means waiting.",
109
+ "steps": [{"tap": "Ceiling", "expect": "the asset list arrives after a count header"}]}
110
+ ```
111
+
112
+ A `stop` names the steps it did not attempt. If the ruling was wrong, re-issue
113
+ them with a corrected `supervise` note. It is off unless asked for, and **not yet
114
+ proven in the field** — measured on a bench, not on a real run.
115
+
116
+ ### Two things that answer a question a screenshot would
117
+
118
+ **Waiting for either of two outcomes.** A login screen *or* a dashboard is a
119
+ disjunction, and asking for it as one intent asks the matcher for something no
120
+ single element answers — one agent spent 120 seconds that way while the login
121
+ screen was already there.
122
+
123
+ ```json
124
+ {"waitFor": {"any": ["Email", "Dashboard"]}}
125
+ ```
126
+
127
+ **A region of the screen, enlarged.** `sim_look` caps at 1024px on the long edge,
128
+ which cannot tell a selected chip from an unselected one. `region` is in points —
129
+ the same coordinates the map prints — and the crop gets the whole budget, so the
130
+ detail per point is the reason to ask:
131
+
132
+ ```json
133
+ {"device": "…", "detail": "high", "region": {"x": 18, "y": 260, "width": 366, "height": 80}}
134
+ ```
135
+
136
+ Reach for it when the question is *selected or not*, *is there a chevron*, *is
137
+ that a validation mark* — the cases text genuinely cannot answer. Not for what a
138
+ field contains or whether a button is enabled; `sim_ui` reports both.
139
+
140
+ ### Anything network-backed: `waitFor`, never `settle`
141
+
142
+ `settle` asks whether the screen has stopped moving. A screen waiting on a
143
+ network call has stopped moving — it is perfectly still and completely empty —
144
+ so `settle` reports success and you act on a list that has not arrived. This
145
+ happened in a real session and cost several calls before the cause was found:
146
+ the settle said settled, the map showed an empty content region, and nothing
147
+ distinguished *empty* from *still loading*.
148
+
149
+ So for anything that has to come over the network — a list, a search result, a
150
+ login, a screen after a submit — assert on the **content you expect**, not on
151
+ stillness:
152
+
153
+ ```json
154
+ {"waitFor": {"value": "Kate Bell", "timeoutMs": 8000}}
155
+ ```
156
+
157
+ `settle` is right for a local transition: a push, a modal, a tab switch. Two
158
+ signals now help you tell the difference without guessing. The map header says
159
+ `STILL LOADING` when the transition classifier can see a load in progress, and
160
+ the `next:` line says the same in words — *an empty-looking region may be a list
161
+ that has not arrived*. Neither is a substitute for asserting on the string you
162
+ are waiting for, which is the only check that knows what "arrived" means.
163
+
164
+ ### A six-step flow in two model turns
165
+
166
+ ```bash
167
+ simframe ui # turn 1: look once
168
+ ```
169
+ ```bash
170
+ simframe do '[ # turn 2: everything else
171
+ {"launch": {"value": "com.example.app"}},
172
+ {"tap": "Sign in"},
173
+ {"type": {"into": "Email", "text": "a@b.com"}},
174
+ {"type": {"into": "Password", "text": "hunter2"}},
175
+ {"assert": {"value": "Sign in", "is": "enabled"}},
176
+ {"tap": "Sign in"},
177
+ {"waitFor": {"value": "Inbox", "timeoutMs": 8000}}
178
+ ]'
179
+ ```
180
+
181
+ Seven steps, one call. The asserts are what make it safe to not look between
182
+ them: if the email did not land, the `assert` halts the flow at that step rather
183
+ than letting the next four run against a screen you were wrong about.
184
+
185
+ ### Recovering from `unexpected-screen`, without starting over
186
+
187
+ ```
188
+ FLOW FAILED — 2/5 steps in 3184ms
189
+ ok [0] tap: tapped "Assets" at 62,835 · settled in 412ms
190
+ FAIL [1] tap: unexpected-screen: expected the screen this action reached 3x
191
+ before, and landed somewhere else
192
+ next: the flow stopped here — this is the moment to think. sim_recall shows how
193
+ you got here; sim_ui re-reads the screen.
194
+ ```
195
+
196
+ That verdict is the tool doing its job: it stopped instead of running three more
197
+ taps on a screen you did not plan for. Recover in two calls, not ten:
198
+
199
+ ```bash
200
+ simframe recall # what happened, as text — not a screenshot
201
+ simframe ui # where you actually are
202
+ ```
203
+
204
+ Then write the *remaining* steps as one new `sim_do`. Do not re-run the steps
205
+ that already succeeded, and do not switch to single taps "to be careful" —
206
+ single taps are the expensive mode, and the asserts are what make batching safe.
207
+
21
208
  ## Read the screen as text, not as an image
22
209
 
23
210
  ```bash
@@ -76,11 +263,13 @@ Steps — every place a control is named accepts a selector:
76
263
  | Act | Check |
77
264
  | --- | --- |
78
265
  | `{"tap": "Save"}` · add `"index"` if a label is ambiguous | `{"assert": {"value": "Saved", "is": "visible"}}` |
79
- | `{"type": {"into": "Name", "text": "..."}}` | `is`: `visible` · `gone` · `enabled` · `disabled` · `value` (with `equals`) |
266
+ | `{"type": {"into": "Name", "text": "..."}}` — drop `into` to type into whatever already has focus | `is`: `visible` · `gone` · `enabled` · `disabled` · `value` (with `equals`) |
80
267
  | `{"paste": {"into": "Notes", "text": "long text"}}` | `{"waitFor": {"value": "Saved", "timeoutMs": 5000}}` |
81
268
  | `{"scroll": "down"}` · `{"scrollTo": "Delete account"}` | `{"settle": {"stableMs": 600}}` |
82
269
  | `{"swipe": {"from": [x,y], "to": [x,y]}}` | `{"pause": 300}` |
83
- | `{"button": "HOME"}` | |
270
+ | `{"button": "HOME"}` — hardware buttons | |
271
+ | `{"key": "return"}` — the **keyboard** return key, which is how a mobile search field submits. Also `escape`, `tab`, `space`, `backspace`, arrows | |
272
+ | `{"clear": "Notes"}` — empty a field. Add `"clear": true` to a `type`/`paste` to **replace** rather than append | |
84
273
  | `{"launch": {"value": "com.example.app", "relaunch": true, "args": ["-uiTest","1"]}}` | |
85
274
  | `{"openUrl": "myapp://path"}` | |
86
275
  | `{"permission": {"value": "photos", "grant": "grant", "bundleId": "com.example.app"}}` | |
@@ -89,12 +278,40 @@ Steps — every place a control is named accepts a selector:
89
278
 
90
279
  | | |
91
280
  | --- | --- |
92
- | `#3` | the number `simframe ui` gave it. Cheapest, and unambiguous. |
93
- | `"Save"` · `the Assets tab` · `back` | resolved by intent — verbs, typos, synonyms, icon-only controls by their common name |
281
+ | `"Save"` · `the Assets tab` · `back` | **start here.** Resolved by intent — verbs, typos, synonyms, icon-only controls by their common name |
282
+ | `#3` | the number `simframe ui` gave it. Cheap and exact, but only within the round trip that numbered it |
94
283
  | `@120,400` | raw point coordinates. Last resort; it cannot tell you it missed. |
95
284
 
285
+ **Prefer a label to a number**, and this order is a correction. Four peer rounds
286
+ in a row reported the same thing: intent resolution worked every time, including
287
+ on labels a string matcher should not have managed, while refs renumbered
288
+ constantly and were safe only inside a single round trip. The table used to lead
289
+ with `#3` and call it "cheapest and unambiguous", which sent every one of them
290
+ down the more brittle path first.
291
+
96
292
  A ref is valid only while that screen is showing. Use one on a different screen
97
- and it refuses rather than tapping whatever now sits at those coordinates.
293
+ and it refuses rather than tapping whatever now sits at those coordinates — and
294
+ the refusal names the label the number was given to, so re-issuing it by label
295
+ costs nothing. A refusal that says the screen *moved* rather than *changed* is
296
+ reporting pixel drift, not a different screen; it says which.
297
+
298
+ ## Lines that mean the tool is unsure, and what to do about each
299
+
300
+ These exist because three peer rounds in one day all reported the same class of
301
+ bug: a component answering confidently when it could not know. Each line below is
302
+ a place that now says so instead. None of them is decoration — if you see one,
303
+ the next call should change.
304
+
305
+ | line | what it means |
306
+ | --- | --- |
307
+ | `[unconfirmed — nothing on this screen reads back the field's contents]` | the text was sent and **nothing verified it landed**. Common on web views, where the accessibility tree carries no field contents. Re-read, or `assert` the value |
308
+ | `a value simframe wrote here is gone: …` | a field simframe filled is on screen and its contents are not. Something cleared it — a reload, a pull-to-refresh, a navigation. Re-fill before continuing |
309
+ | `N element(s) on this screen overlap and disagree about what is there` | a sheet or overlay is probably covering the screen behind it, and some elements belong to the layer underneath |
310
+ | `WARNING: this image is Ns older than the screen state` | the picture is very likely not what is on the device. Use `sim_ui`, which is read live |
311
+ | `#N cannot be trusted here — … layout distance D, tolerance T` | pixel drift, **not** a different screen. The identity may be unchanged; re-issue by label |
312
+ | `#N cannot be trusted here — … this is a different screen` | the screen really did change. Read it again |
313
+ | `autoSettle was off, so the map below was read without waiting` | the map may describe the screen *before* the last action landed |
314
+ | `"iPhone 17 Pro" names more than one booted device` | a name cannot identify which device answered. Pass `device` with a UDID |
98
315
 
99
316
  ## Every step is verified, and the verdict means something
100
317
 
@@ -151,6 +368,21 @@ simframe do /tmp/flow.json --json | jq '.results[] | select(.ok==false)'
151
368
  3. `simframe do` — act, in a batch, with asserts inside the batch.
152
369
  4. `simframe frame` / `sim_look` — pixels. Only for a question about pixels.
153
370
 
371
+ A screenshot is about 1600 tokens and it is the most expensive call here. Before
372
+ reaching for one, check it is not a question the text already answers: a map row
373
+ carries the element's **contents** (`= Fryer 3`) and its **state** (`disabled`),
374
+ and the flow's own verdict already said whether the action worked. Those three
375
+ account for nearly every screenshot taken in the session that was measured —
376
+ 28 of 62 calls returned an image, about a third of that session's entire token
377
+ cost.
378
+
379
+ ## One goal per session
380
+
381
+ A session gets slower with every turn: more context to carry, and later turns
382
+ run measurably longer than early ones. Where you can, give one test goal its own
383
+ session and finish it. `sim_do` is what keeps a session short — a flow that runs
384
+ as one call adds one exchange to the context instead of twelve.
385
+
154
386
  ## When something is wrong with simframe itself
155
387
 
156
388
  ```bash