simframe 0.10.0 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +137 -2
- package/data/vocabulary/en.json +148 -0
- package/native/ocr.swift +13 -1
- package/native/rank.swift +87 -0
- package/native/simframed/Sources/PrivateAPI/CoreSimulatorPlatform.swift +43 -3
- package/native/simframed/Sources/PrivateAPI/PrivateAPI.swift +27 -0
- package/native/simframed/Sources/PrivateAPI/StubPlatform.swift +4 -0
- package/native/simframed/Sources/simframed/main.swift +13 -1
- package/native/supervise.swift +181 -0
- package/package.json +4 -1
- package/scripts/check-package.mjs +22 -2
- package/scripts/ci-memory.mjs +104 -20
- package/scripts/eval-perception.mjs +33 -0
- package/scripts/phase17-corpus.mjs +176 -0
- package/skills/simframe/SKILL.md +237 -5
- package/src/actions.js +1692 -44
- package/src/cli.js +131 -9
- package/src/control.js +1 -0
- package/src/fingerprint.js +19 -1
- package/src/graph.js +89 -7
- package/src/index.js +211 -14
- package/src/input.js +66 -3
- package/src/localhelper.js +155 -0
- package/src/matching.js +64 -1
- package/src/mcp.js +319 -27
- package/src/metrics.js +32 -3
- package/src/ocr.js +18 -1
- package/src/planner.js +195 -0
- package/src/platform/android.js +2 -1
- package/src/platform/ios.js +2 -1
- package/src/png.js +26 -0
- package/src/refs.js +51 -8
- package/src/regions.js +110 -1
- package/src/screenmap.js +89 -9
- package/src/supervisor.js +117 -0
- package/src/view.js +335 -7
- package/src/vocabulary.js +134 -0
- package/src/wrote.js +136 -0
package/skills/simframe/SKILL.md
CHANGED
|
@@ -18,6 +18,193 @@ and CV alone. Tapping by label works; screen recognition is thinner, so prefer
|
|
|
18
18
|
naming a device explicitly and re-reading the screen after a step you are unsure
|
|
19
19
|
about.
|
|
20
20
|
|
|
21
|
+
## The protocol: plan once, execute once, think only when told to
|
|
22
|
+
|
|
23
|
+
The expensive thing in a simulator session is not the tapping. It is you —
|
|
24
|
+
observe, think, tap, observe, think. Measured over a real session: 62 tool
|
|
25
|
+
calls for 179 steps, and 48 of those calls were three steps or fewer. A
|
|
26
|
+
twelve-step flow arrived as five calls, and every boundary between them was a
|
|
27
|
+
think that nothing had asked for.
|
|
28
|
+
|
|
29
|
+
So the protocol is four steps, and step 4 is the one that saves the time.
|
|
30
|
+
|
|
31
|
+
1. **State the goal, then read the screen once.** `simframe ui`, or `sim_find`
|
|
32
|
+
if you doubt a selector will resolve. Once — not per step.
|
|
33
|
+
2. **Write the entire flow as one `sim_do`**, with an assert after each step
|
|
34
|
+
whose success you would otherwise have checked by looking.
|
|
35
|
+
3. **Run it. Read only the verdict line per step.**
|
|
36
|
+
4. **Think again only when the result tells you to.** Every action result ends
|
|
37
|
+
with a `next:` line the daemon computed locally, and it says which case you
|
|
38
|
+
are in:
|
|
39
|
+
|
|
40
|
+
- `next: settled; screen known (…); N elements; nothing ambiguous — chain the
|
|
41
|
+
next steps in one sim_do without looking again` → **do not look. Act.**
|
|
42
|
+
- `next: new screen, nothing predicted here yet` → read it before acting on a
|
|
43
|
+
label you have not seen on it.
|
|
44
|
+
- `next: N labels repeat on this screen` → address those by `#ref`.
|
|
45
|
+
- `next: the flow stopped here` → this is the moment to think.
|
|
46
|
+
- `worked here before: tap "…" (7x)` → the graph's own vocabulary for this
|
|
47
|
+
screen, most-used first. It is evidence for writing a chain, not an
|
|
48
|
+
instruction: the most-tapped control on a screen is often what earlier runs
|
|
49
|
+
used to *back out*.
|
|
50
|
+
- `memory disagrees with this screen: N remembered controls not present` →
|
|
51
|
+
two screens share one fingerprint. Trust the element list, not the graph,
|
|
52
|
+
and re-read before anything irreversible.
|
|
53
|
+
|
|
54
|
+
**Do not** narrate each step, re-read the screen after every action, or use
|
|
55
|
+
extended thinking inside a flow. The flow is already planned; executing it is
|
|
56
|
+
not a decision.
|
|
57
|
+
|
|
58
|
+
### Recovering without a round trip
|
|
59
|
+
|
|
60
|
+
Three things let a batch survive a problem instead of handing it back. Every one
|
|
61
|
+
of them exists because a real run lost a call to the thing it prevents.
|
|
62
|
+
|
|
63
|
+
**Fallback selectors.** `{"tap": "Save", "or": ["Done", "Confirm"]}` — tried
|
|
64
|
+
locally, in order, and only an exhausted list reaches you. Eligible after a
|
|
65
|
+
selector that did not *resolve*, and nothing else: retrying from a screen you did
|
|
66
|
+
not expect to be on is a second guess, not a retry. simframe refuses to
|
|
67
|
+
substitute a label that looks destructive even when you list it.
|
|
68
|
+
|
|
69
|
+
They are a cure for *"I named it wrong"*, and most real failures are *"it is not
|
|
70
|
+
there yet"* — so when everything in an `or` chain misses and the screen has only
|
|
71
|
+
just stopped moving, the failure says so. Reach for `waitFor` there, not for more
|
|
72
|
+
labels.
|
|
73
|
+
|
|
74
|
+
**`{"seek": "change username", "budget": 6}`** looks for something that is not on
|
|
75
|
+
this screen: it opens containers, checks, and comes back, depth first.
|
|
76
|
+
|
|
77
|
+
It **acts** — opening a door changes state — and it refuses to open anything that
|
|
78
|
+
commits, abandons or answers. It once opened `CANCEL` and pressed *"YES, THIS
|
|
79
|
+
FIXED MY PROBLEM"*, which is why that sentence is here. It does not tap the
|
|
80
|
+
target: it leaves you on the screen where the target resolves, and returns to
|
|
81
|
+
where it started if it fails. **Do not point it into a flow whose progress you
|
|
82
|
+
cannot afford to lose.**
|
|
83
|
+
|
|
84
|
+
**`sweep` — read and fill a long screen.** A viewport is the only honest unit:
|
|
85
|
+
the tree publishes what is rendered, so a form taller than the screen is knowable
|
|
86
|
+
only in pieces.
|
|
87
|
+
|
|
88
|
+
```json
|
|
89
|
+
{"sweep": "all", "fill": {"Last Name": "Asadi", "Email": "a@b.c", "Comment": "…"}}
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
It goes to the top, then reads section by section to the bottom, filling each
|
|
93
|
+
field **while it is on screen** — which beats finding a field and then trying to
|
|
94
|
+
scroll back to it, because one gesture travels a non-deterministic distance.
|
|
95
|
+
`{"from": "here"}` sweeps down from where you are. `{"sweep": "<text>"}` stops as
|
|
96
|
+
soon as it finds that text. It reports which section each element was in, what it
|
|
97
|
+
filled, and what it never found at any scroll position.
|
|
98
|
+
|
|
99
|
+
Prefer it to `scrollTo` on any form or long list. `scrollTo` hunts one label and
|
|
100
|
+
cannot work when the label is not yet rendered; a sweep covers the screen.
|
|
101
|
+
|
|
102
|
+
**The local supervisor**, when one is enabled, decides whether a failed step
|
|
103
|
+
should `wait`, `retry` or `stop` — before the failure reaches you. It knows
|
|
104
|
+
nothing about the app and you do, so brief it from the plan:
|
|
105
|
+
|
|
106
|
+
```json
|
|
107
|
+
{"supervisor": "apple",
|
|
108
|
+
"supervise": "Lists here render a count header before their rows, so a missing row usually means waiting.",
|
|
109
|
+
"steps": [{"tap": "Ceiling", "expect": "the asset list arrives after a count header"}]}
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
A `stop` names the steps it did not attempt. If the ruling was wrong, re-issue
|
|
113
|
+
them with a corrected `supervise` note. It is off unless asked for, and **not yet
|
|
114
|
+
proven in the field** — measured on a bench, not on a real run.
|
|
115
|
+
|
|
116
|
+
### Two things that answer a question a screenshot would
|
|
117
|
+
|
|
118
|
+
**Waiting for either of two outcomes.** A login screen *or* a dashboard is a
|
|
119
|
+
disjunction, and asking for it as one intent asks the matcher for something no
|
|
120
|
+
single element answers — one agent spent 120 seconds that way while the login
|
|
121
|
+
screen was already there.
|
|
122
|
+
|
|
123
|
+
```json
|
|
124
|
+
{"waitFor": {"any": ["Email", "Dashboard"]}}
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
**A region of the screen, enlarged.** `sim_look` caps at 1024px on the long edge,
|
|
128
|
+
which cannot tell a selected chip from an unselected one. `region` is in points —
|
|
129
|
+
the same coordinates the map prints — and the crop gets the whole budget, so the
|
|
130
|
+
detail per point is the reason to ask:
|
|
131
|
+
|
|
132
|
+
```json
|
|
133
|
+
{"device": "…", "detail": "high", "region": {"x": 18, "y": 260, "width": 366, "height": 80}}
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
Reach for it when the question is *selected or not*, *is there a chevron*, *is
|
|
137
|
+
that a validation mark* — the cases text genuinely cannot answer. Not for what a
|
|
138
|
+
field contains or whether a button is enabled; `sim_ui` reports both.
|
|
139
|
+
|
|
140
|
+
### Anything network-backed: `waitFor`, never `settle`
|
|
141
|
+
|
|
142
|
+
`settle` asks whether the screen has stopped moving. A screen waiting on a
|
|
143
|
+
network call has stopped moving — it is perfectly still and completely empty —
|
|
144
|
+
so `settle` reports success and you act on a list that has not arrived. This
|
|
145
|
+
happened in a real session and cost several calls before the cause was found:
|
|
146
|
+
the settle said settled, the map showed an empty content region, and nothing
|
|
147
|
+
distinguished *empty* from *still loading*.
|
|
148
|
+
|
|
149
|
+
So for anything that has to come over the network — a list, a search result, a
|
|
150
|
+
login, a screen after a submit — assert on the **content you expect**, not on
|
|
151
|
+
stillness:
|
|
152
|
+
|
|
153
|
+
```json
|
|
154
|
+
{"waitFor": {"value": "Kate Bell", "timeoutMs": 8000}}
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
`settle` is right for a local transition: a push, a modal, a tab switch. Two
|
|
158
|
+
signals now help you tell the difference without guessing. The map header says
|
|
159
|
+
`STILL LOADING` when the transition classifier can see a load in progress, and
|
|
160
|
+
the `next:` line says the same in words — *an empty-looking region may be a list
|
|
161
|
+
that has not arrived*. Neither is a substitute for asserting on the string you
|
|
162
|
+
are waiting for, which is the only check that knows what "arrived" means.
|
|
163
|
+
|
|
164
|
+
### A six-step flow in two model turns
|
|
165
|
+
|
|
166
|
+
```bash
|
|
167
|
+
simframe ui # turn 1: look once
|
|
168
|
+
```
|
|
169
|
+
```bash
|
|
170
|
+
simframe do '[ # turn 2: everything else
|
|
171
|
+
{"launch": {"value": "com.example.app"}},
|
|
172
|
+
{"tap": "Sign in"},
|
|
173
|
+
{"type": {"into": "Email", "text": "a@b.com"}},
|
|
174
|
+
{"type": {"into": "Password", "text": "hunter2"}},
|
|
175
|
+
{"assert": {"value": "Sign in", "is": "enabled"}},
|
|
176
|
+
{"tap": "Sign in"},
|
|
177
|
+
{"waitFor": {"value": "Inbox", "timeoutMs": 8000}}
|
|
178
|
+
]'
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
Seven steps, one call. The asserts are what make it safe to not look between
|
|
182
|
+
them: if the email did not land, the `assert` halts the flow at that step rather
|
|
183
|
+
than letting the next four run against a screen you were wrong about.
|
|
184
|
+
|
|
185
|
+
### Recovering from `unexpected-screen`, without starting over
|
|
186
|
+
|
|
187
|
+
```
|
|
188
|
+
FLOW FAILED — 2/5 steps in 3184ms
|
|
189
|
+
ok [0] tap: tapped "Assets" at 62,835 · settled in 412ms
|
|
190
|
+
FAIL [1] tap: unexpected-screen: expected the screen this action reached 3x
|
|
191
|
+
before, and landed somewhere else
|
|
192
|
+
next: the flow stopped here — this is the moment to think. sim_recall shows how
|
|
193
|
+
you got here; sim_ui re-reads the screen.
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
That verdict is the tool doing its job: it stopped instead of running three more
|
|
197
|
+
taps on a screen you did not plan for. Recover in two calls, not ten:
|
|
198
|
+
|
|
199
|
+
```bash
|
|
200
|
+
simframe recall # what happened, as text — not a screenshot
|
|
201
|
+
simframe ui # where you actually are
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
Then write the *remaining* steps as one new `sim_do`. Do not re-run the steps
|
|
205
|
+
that already succeeded, and do not switch to single taps "to be careful" —
|
|
206
|
+
single taps are the expensive mode, and the asserts are what make batching safe.
|
|
207
|
+
|
|
21
208
|
## Read the screen as text, not as an image
|
|
22
209
|
|
|
23
210
|
```bash
|
|
@@ -76,11 +263,13 @@ Steps — every place a control is named accepts a selector:
|
|
|
76
263
|
| Act | Check |
|
|
77
264
|
| --- | --- |
|
|
78
265
|
| `{"tap": "Save"}` · add `"index"` if a label is ambiguous | `{"assert": {"value": "Saved", "is": "visible"}}` |
|
|
79
|
-
| `{"type": {"into": "Name", "text": "..."}}` | `is`: `visible` · `gone` · `enabled` · `disabled` · `value` (with `equals`) |
|
|
266
|
+
| `{"type": {"into": "Name", "text": "..."}}` — drop `into` to type into whatever already has focus | `is`: `visible` · `gone` · `enabled` · `disabled` · `value` (with `equals`) |
|
|
80
267
|
| `{"paste": {"into": "Notes", "text": "long text"}}` | `{"waitFor": {"value": "Saved", "timeoutMs": 5000}}` |
|
|
81
268
|
| `{"scroll": "down"}` · `{"scrollTo": "Delete account"}` | `{"settle": {"stableMs": 600}}` |
|
|
82
269
|
| `{"swipe": {"from": [x,y], "to": [x,y]}}` | `{"pause": 300}` |
|
|
83
|
-
| `{"button": "HOME"}` | |
|
|
270
|
+
| `{"button": "HOME"}` — hardware buttons | |
|
|
271
|
+
| `{"key": "return"}` — the **keyboard** return key, which is how a mobile search field submits. Also `escape`, `tab`, `space`, `backspace`, arrows | |
|
|
272
|
+
| `{"clear": "Notes"}` — empty a field. Add `"clear": true` to a `type`/`paste` to **replace** rather than append | |
|
|
84
273
|
| `{"launch": {"value": "com.example.app", "relaunch": true, "args": ["-uiTest","1"]}}` | |
|
|
85
274
|
| `{"openUrl": "myapp://path"}` | |
|
|
86
275
|
| `{"permission": {"value": "photos", "grant": "grant", "bundleId": "com.example.app"}}` | |
|
|
@@ -89,12 +278,40 @@ Steps — every place a control is named accepts a selector:
|
|
|
89
278
|
|
|
90
279
|
| | |
|
|
91
280
|
| --- | --- |
|
|
92
|
-
|
|
|
93
|
-
| `
|
|
281
|
+
| `"Save"` · `the Assets tab` · `back` | **start here.** Resolved by intent — verbs, typos, synonyms, icon-only controls by their common name |
|
|
282
|
+
| `#3` | the number `simframe ui` gave it. Cheap and exact, but only within the round trip that numbered it |
|
|
94
283
|
| `@120,400` | raw point coordinates. Last resort; it cannot tell you it missed. |
|
|
95
284
|
|
|
285
|
+
**Prefer a label to a number**, and this order is a correction. Four peer rounds
|
|
286
|
+
in a row reported the same thing: intent resolution worked every time, including
|
|
287
|
+
on labels a string matcher should not have managed, while refs renumbered
|
|
288
|
+
constantly and were safe only inside a single round trip. The table used to lead
|
|
289
|
+
with `#3` and call it "cheapest and unambiguous", which sent every one of them
|
|
290
|
+
down the more brittle path first.
|
|
291
|
+
|
|
96
292
|
A ref is valid only while that screen is showing. Use one on a different screen
|
|
97
|
-
and it refuses rather than tapping whatever now sits at those coordinates
|
|
293
|
+
and it refuses rather than tapping whatever now sits at those coordinates — and
|
|
294
|
+
the refusal names the label the number was given to, so re-issuing it by label
|
|
295
|
+
costs nothing. A refusal that says the screen *moved* rather than *changed* is
|
|
296
|
+
reporting pixel drift, not a different screen; it says which.
|
|
297
|
+
|
|
298
|
+
## Lines that mean the tool is unsure, and what to do about each
|
|
299
|
+
|
|
300
|
+
These exist because three peer rounds in one day all reported the same class of
|
|
301
|
+
bug: a component answering confidently when it could not know. Each line below is
|
|
302
|
+
a place that now says so instead. None of them is decoration — if you see one,
|
|
303
|
+
the next call should change.
|
|
304
|
+
|
|
305
|
+
| line | what it means |
|
|
306
|
+
| --- | --- |
|
|
307
|
+
| `[unconfirmed — nothing on this screen reads back the field's contents]` | the text was sent and **nothing verified it landed**. Common on web views, where the accessibility tree carries no field contents. Re-read, or `assert` the value |
|
|
308
|
+
| `a value simframe wrote here is gone: …` | a field simframe filled is on screen and its contents are not. Something cleared it — a reload, a pull-to-refresh, a navigation. Re-fill before continuing |
|
|
309
|
+
| `N element(s) on this screen overlap and disagree about what is there` | a sheet or overlay is probably covering the screen behind it, and some elements belong to the layer underneath |
|
|
310
|
+
| `WARNING: this image is Ns older than the screen state` | the picture is very likely not what is on the device. Use `sim_ui`, which is read live |
|
|
311
|
+
| `#N cannot be trusted here — … layout distance D, tolerance T` | pixel drift, **not** a different screen. The identity may be unchanged; re-issue by label |
|
|
312
|
+
| `#N cannot be trusted here — … this is a different screen` | the screen really did change. Read it again |
|
|
313
|
+
| `autoSettle was off, so the map below was read without waiting` | the map may describe the screen *before* the last action landed |
|
|
314
|
+
| `"iPhone 17 Pro" names more than one booted device` | a name cannot identify which device answered. Pass `device` with a UDID |
|
|
98
315
|
|
|
99
316
|
## Every step is verified, and the verdict means something
|
|
100
317
|
|
|
@@ -151,6 +368,21 @@ simframe do /tmp/flow.json --json | jq '.results[] | select(.ok==false)'
|
|
|
151
368
|
3. `simframe do` — act, in a batch, with asserts inside the batch.
|
|
152
369
|
4. `simframe frame` / `sim_look` — pixels. Only for a question about pixels.
|
|
153
370
|
|
|
371
|
+
A screenshot is about 1600 tokens and it is the most expensive call here. Before
|
|
372
|
+
reaching for one, check it is not a question the text already answers: a map row
|
|
373
|
+
carries the element's **contents** (`= Fryer 3`) and its **state** (`disabled`),
|
|
374
|
+
and the flow's own verdict already said whether the action worked. Those three
|
|
375
|
+
account for nearly every screenshot taken in the session that was measured —
|
|
376
|
+
28 of 62 calls returned an image, about a third of that session's entire token
|
|
377
|
+
cost.
|
|
378
|
+
|
|
379
|
+
## One goal per session
|
|
380
|
+
|
|
381
|
+
A session gets slower with every turn: more context to carry, and later turns
|
|
382
|
+
run measurably longer than early ones. Where you can, give one test goal its own
|
|
383
|
+
session and finish it. `sim_do` is what keeps a session short — a flow that runs
|
|
384
|
+
as one call adds one exchange to the context instead of twelve.
|
|
385
|
+
|
|
154
386
|
## When something is wrong with simframe itself
|
|
155
387
|
|
|
156
388
|
```bash
|