simframe 0.6.2 → 0.7.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +112 -8
- package/native/simframed/Sources/SimframeCore/CaptureRecovery.swift +34 -0
- package/native/simframed/Sources/SimframeCore/FrameStore.swift +20 -0
- package/native/simframed/Sources/simframed/main.swift +30 -0
- package/native/simframed/Tests/SimframeCoreTests/HashingTests.swift +39 -0
- package/package.json +4 -2
- package/scripts/analyse-fingerprint.mjs +149 -0
- package/scripts/bench-flow.mjs +1 -1
- package/scripts/check-package.mjs +15 -0
- package/scripts/ci-memory.mjs +12 -0
- package/scripts/eval-ax-tier.mjs +191 -0
- package/scripts/eval-fingerprint.mjs +161 -6
- package/skills/simframe/SKILL.md +10 -3
- package/src/actions.js +64 -14
- package/src/cli.js +125 -32
- package/src/daemon.js +42 -1
- package/src/engine.js +19 -6
- package/src/fingerprint.js +29 -1
- package/src/index.js +92 -13
- package/src/input.js +93 -0
- package/src/mcp.js +6 -6
- package/src/navigate.js +8 -1
- package/src/platform/android.js +997 -0
- package/src/platform/host.js +15 -0
- package/src/platform/index.js +263 -0
- package/src/{simctl.js → platform/ios.js} +111 -18
- package/src/store.js +32 -0
package/README.md
CHANGED
|
@@ -4,11 +4,11 @@
|
|
|
4
4
|
[](https://www.npmjs.com/package/simframe)
|
|
5
5
|
[](./LICENSE)
|
|
6
6
|
|
|
7
|
-
**Eyes, hands and memory for an agent driving the iOS Simulator.**
|
|
7
|
+
**Eyes, hands and memory for an agent driving the iOS Simulator or an Android emulator.**
|
|
8
8
|
|
|
9
9
|
[Website](https://lvlrsajjad.github.io/simframe/) · [npm](https://www.npmjs.com/package/simframe)
|
|
10
10
|
|
|
11
|
-
An agent driving
|
|
11
|
+
An agent driving a simulator is slow for three reasons, and only the first
|
|
12
12
|
one is obvious:
|
|
13
13
|
|
|
14
14
|
1. **Every look is a wait.** `simctl io screenshot` costs ~130 ms of blocking
|
|
@@ -129,6 +129,87 @@ breaks the host-side path.
|
|
|
129
129
|
brew tap facebook/fb && brew install idb-companion && pipx install fb-idb
|
|
130
130
|
```
|
|
131
131
|
|
|
132
|
+
## Android
|
|
133
|
+
|
|
134
|
+
simframe's second backend drives an Android emulator with the same commands, the
|
|
135
|
+
same screen map and the same memory as a simulator. Everything above the
|
|
136
|
+
platform boundary — the frame store, settle, the structural fingerprint, the
|
|
137
|
+
screen map, refs and the transition graph — runs on it unmodified, because the
|
|
138
|
+
boundary hands it frames and nothing above it knows what a simulator is.
|
|
139
|
+
|
|
140
|
+
| Capability | Android | How |
|
|
141
|
+
| --- | --- | --- |
|
|
142
|
+
| Watch the screen, wait, recall | yes | frames at **41 ms** through the emulator console, host-side — no adb in the capture path |
|
|
143
|
+
| Read labels + coordinates from pixels | yes | the same Vision OCR + CV, off the same PNG |
|
|
144
|
+
| Screen map, refs, screen memory, the graph | yes | unchanged above the boundary |
|
|
145
|
+
| Tap, type, swipe, keys | yes | the console's `event mouse` as a real down/move/up, `event text` for characters, `input keyevent` for keys |
|
|
146
|
+
| Clipboard, and `paste` into a field | yes | the emulator's gRPC `setClipboard`, over `node:http2`, no dependency, then `KEYCODE_PASTE` to deliver it |
|
|
147
|
+
| List/resolve devices, launch, terminate, open a URL, permissions | yes | `adb`, with the permission state read back off the device |
|
|
148
|
+
| Accessibility tree | **not available (OCR + CV only)** | `uiautomator dump` costs **2,012 ms** a read, against 45 ms for the iOS tree. See [`docs/DEFERRED.md`](docs/DEFERRED.md) |
|
|
149
|
+
|
|
150
|
+
```bash
|
|
151
|
+
# an emulator is found the same way a simulator is
|
|
152
|
+
simframe devices # ● Small_Phone_API_36 Android 16 (API 36) emulator-5554
|
|
153
|
+
simframe ui --device=emulator-5554
|
|
154
|
+
simframe do --device=emulator-5554 flow.json
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
A host with a booted simulator **and** a booted emulator has no default, and
|
|
158
|
+
simframe will not pick one for you: preferring iOS because it came first would
|
|
159
|
+
tap a simulator while you were driving an emulator, and acting on the wrong
|
|
160
|
+
device is worse than refusing. So a command with no device names both and stops.
|
|
161
|
+
`--device` answers it per command; `SIMFRAME_DEVICE` answers it per shell:
|
|
162
|
+
|
|
163
|
+
```bash
|
|
164
|
+
export SIMFRAME_DEVICE=emulator-5554
|
|
165
|
+
simframe ui # the emulator, without saying so every time
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
The tree is a deliberate omission, not an oversight. Making it fast needs a
|
|
169
|
+
resident instrumentation APK on the device — the shape uiautomator2, Maestro and
|
|
170
|
+
Appium all converged on — and that would be simframe's first runtime artifact
|
|
171
|
+
installed onto your device. The perception ladder was built so a missing tier
|
|
172
|
+
degrades rather than fails, and this is exactly that case: OCR and CV yield
|
|
173
|
+
labels and coordinates on Android today, and a tap by label works without a tree
|
|
174
|
+
at all. `simframe doctor` reports the tier as `optional` with that number, so
|
|
175
|
+
the gap is visible rather than silent, and the criteria for revisiting it are in
|
|
176
|
+
`docs/DEFERRED.md` under **Phase 8b**.
|
|
177
|
+
|
|
178
|
+
**What the missing tree costs, measured rather than hand-waved.** Screen
|
|
179
|
+
*identity* is weaker on Android than on iOS, and specifically so. Tokens per
|
|
180
|
+
screen, and where they come from:
|
|
181
|
+
|
|
182
|
+
| Screen | Tokens | Regions | Roles | Chrome labels |
|
|
183
|
+
| --- | --- | --- | --- | --- |
|
|
184
|
+
| launcher | **1** | nav-bar 1 | text 1 | 0 |
|
|
185
|
+
| Settings root | 9 | content 9 | text 9 | 0 |
|
|
186
|
+
| example.com in Chrome | 6 | content 3, nav-bar 3 | text 6 | 0 |
|
|
187
|
+
|
|
188
|
+
Every token has role `text`, because without a tree nothing infers a button from
|
|
189
|
+
a rectangle reliably enough to say so, and no screen here carries a chrome label
|
|
190
|
+
at all. So on Android a screen is recognised by the geometry of its text, which
|
|
191
|
+
is thinner and noisier than the iOS mix of roles, chrome labels and geometry.
|
|
192
|
+
Flows still work; screen *memory* is doing more guessing, and that is the honest
|
|
193
|
+
cost of the tier being absent.
|
|
194
|
+
|
|
195
|
+
It is also why the obvious fix for the iOS drift — dropping content-region text
|
|
196
|
+
out of identity, which would be a strict improvement there — is not available:
|
|
197
|
+
it would leave Settings' root with zero tokens, and zero tokens is no identity
|
|
198
|
+
at all. See [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
199
|
+
|
|
200
|
+
That last column read **3** before this measurement changed it. Chrome's address
|
|
201
|
+
bar is chrome by every structural test there is, so the screen's identity
|
|
202
|
+
contained `"== example.com"` — a URL, meaning the same browser on a different
|
|
203
|
+
page was a different screen and every learned route through it broke on
|
|
204
|
+
navigation — plus `":"` and `"+"`, which are OCR reading punctuation off icons.
|
|
205
|
+
A chrome label now has to be a name: two letters at minimum, and not an
|
|
206
|
+
address.
|
|
207
|
+
|
|
208
|
+
The emulator's own gRPC surface was checked for anything tree-shaped and has
|
|
209
|
+
nothing: 43 RPCs for sensors, input, screenshots and VM state, and no notion of
|
|
210
|
+
a view. That question is settled, not open. The same surface is what carries the
|
|
211
|
+
clipboard.
|
|
212
|
+
|
|
132
213
|
## The tools
|
|
133
214
|
|
|
134
215
|
Read first, act in batches, and look at pixels only when the question is about
|
|
@@ -397,7 +478,10 @@ an Xcode upgrade that moves something is a bounded fix rather than an
|
|
|
397
478
|
archaeology project. If a layer breaks, simframe degrades to the layer below
|
|
398
479
|
and `doctor` says which.
|
|
399
480
|
|
|
400
|
-
Run `simframe start --engine=
|
|
481
|
+
Run `simframe start --engine=screenshot` to use the one-frame-at-a-time loop
|
|
482
|
+
instead — it is also the only capture engine on Android, where it reaches the
|
|
483
|
+
emulator console rather than any simulator tool. `--engine=simctl` is still
|
|
484
|
+
accepted as the name that loop used to have.
|
|
401
485
|
|
|
402
486
|
- **Files are the IPC for reads.** The daemon renames completed frames into
|
|
403
487
|
place; readers just read them. A rename is atomic, so a reader can never see a
|
|
@@ -444,8 +528,14 @@ simframe frame --out=now.png # newest frame, native resolution, to a file
|
|
|
444
528
|
simframe strip --count=6 # contact sheet, for an animation
|
|
445
529
|
simframe doctor --strict # any degraded layer is a non-zero exit
|
|
446
530
|
simframe start / status / stop [--force] / devices
|
|
531
|
+
simframe ui --device=emulator-5554 # or export SIMFRAME_DEVICE once
|
|
447
532
|
```
|
|
448
533
|
|
|
534
|
+
`--device` takes `--device=X` and `--device X` alike. It used to take only the
|
|
535
|
+
first: the space form set the flag to `true` and then resolved a device named
|
|
536
|
+
"true", which is a poor answer to a flag `doctor`'s own advice tells you to
|
|
537
|
+
type.
|
|
538
|
+
|
|
449
539
|
### The Claude Code skill
|
|
450
540
|
|
|
451
541
|
[`skills/simframe/SKILL.md`](skills/simframe/SKILL.md) teaches the CLI path
|
|
@@ -488,6 +578,13 @@ So every downgrade now announces itself:
|
|
|
488
578
|
is deliberate: `WARN` means this machine could be doing better and silently is
|
|
489
579
|
not, which is the failure worth shouting about. An optional fallback missing on
|
|
490
580
|
a fresh machine has not degraded from anything, and `--strict` ignores it.
|
|
581
|
+
- A device whose capture has **wedged** says so: `capture: stalled — the display
|
|
582
|
+
surface has been unreadable for 62s; 3 re-attaches did not help; only
|
|
583
|
+
restarting the device is known to cure it`. This is a different thing from a
|
|
584
|
+
still screen, and it used to look identical, because a damage-driven engine
|
|
585
|
+
produces no frames for either. An agent told "nothing changed" keeps tapping;
|
|
586
|
+
one told the simulator is wedged stops. simframe reports it and does not
|
|
587
|
+
restart your device.
|
|
491
588
|
- `--strict`, or `SIMFRAME_STRICT=1`, turns any downgrade into a non-zero exit.
|
|
492
589
|
CI runs strict, so a release cannot ship in the state that shipped twice.
|
|
493
590
|
|
|
@@ -511,12 +608,15 @@ said a word — the exact failure shape, found by the thing built to catch it.
|
|
|
511
608
|
|
|
512
609
|
## Limitations
|
|
513
610
|
|
|
514
|
-
- Simulators only. Neither the framebuffer nor `simctl`
|
|
515
|
-
device.
|
|
611
|
+
- Simulators and Android emulators only. Neither the framebuffer nor `simctl`
|
|
612
|
+
nor the emulator console can reach a physical device.
|
|
613
|
+
- Android has no accessibility tree, so its screen identity rests on the
|
|
614
|
+
geometry of OCR'd text: thinner and noisier than iOS's. See
|
|
615
|
+
[Android](#android) above.
|
|
516
616
|
- The daemon depends on private frameworks. They are stable enough to build on —
|
|
517
617
|
capture and accessibility survived the iOS 26 transition — but an Xcode
|
|
518
618
|
upgrade can move a symbol. `doctor` reports each layer separately so a break
|
|
519
|
-
is visible rather than mysterious, and `--engine=
|
|
619
|
+
is visible rather than mysterious, and `--engine=screenshot` still works.
|
|
520
620
|
- Hardware buttons: only `home` is implemented. The other Indigo codes are
|
|
521
621
|
unverified, and a wrong one can crash `backboardd` or lock the device, so they
|
|
522
622
|
return an error rather than a guess.
|
|
@@ -532,9 +632,13 @@ said a word — the exact failure shape, found by the thing built to catch it.
|
|
|
532
632
|
|
|
533
633
|
## Roadmap
|
|
534
634
|
|
|
535
|
-
- **Android, as a second backend.** Everything above the platform boundary is
|
|
536
|
-
already platform-agnostic; nothing above it imports a simulator framework.
|
|
537
635
|
- **Extend the confirm vocabulary beyond English.**
|
|
636
|
+
- **Region bands from clustering**, replacing the positional bands. They have
|
|
637
|
+
produced three bugs in three phases, and on Android they put a URL bar in the
|
|
638
|
+
nav bar and its URL into the screen's identity.
|
|
639
|
+
- **Phase 8b, conditionally:** an instrumentation APK for the Android
|
|
640
|
+
accessibility tree, with the criteria for doing it stated in
|
|
641
|
+
`docs/DEFERRED.md` rather than left to enthusiasm.
|
|
538
642
|
|
|
539
643
|
## Releasing
|
|
540
644
|
|
|
@@ -18,15 +18,48 @@ public struct CaptureRecovery {
|
|
|
18
18
|
/// watches a dead capture loop and wonders.
|
|
19
19
|
public static let reattachAfterFailures = 6
|
|
20
20
|
|
|
21
|
+
/// When re-resolving the port has demonstrably not helped.
|
|
22
|
+
///
|
|
23
|
+
/// Re-resolving *succeeds* in the pathology this exists for: the call
|
|
24
|
+
/// returns a fresh descriptor, the callback re-arms, and every read still
|
|
25
|
+
/// fails. Because a successful re-resolve resets the failure count, that
|
|
26
|
+
/// state loops — six failures, re-resolve, six failures — and no count of
|
|
27
|
+
/// consecutive failures ever grows large enough to notice it. Observed
|
|
28
|
+
/// three times in one afternoon on a simulator driven hard for ten minutes;
|
|
29
|
+
/// only restarting the device cured it.
|
|
30
|
+
///
|
|
31
|
+
/// So the signal is re-resolves, not failures. Two of them means the port
|
|
32
|
+
/// was not the problem.
|
|
33
|
+
public static let stalledAfterReattaches = 2
|
|
34
|
+
|
|
35
|
+
/// And the other way it goes wrong: re-resolving itself failing, where the
|
|
36
|
+
/// failure count does keep growing because nothing resets it.
|
|
37
|
+
public static let stalledAfterFailures = reattachAfterFailures * 3
|
|
38
|
+
|
|
21
39
|
public private(set) var consecutiveFailures = 0
|
|
40
|
+
/// Successful re-resolves since the last real frame.
|
|
41
|
+
public private(set) var reattaches = 0
|
|
22
42
|
private let threshold: Int
|
|
23
43
|
|
|
24
44
|
public init(threshold: Int = CaptureRecovery.reattachAfterFailures) {
|
|
25
45
|
self.threshold = threshold
|
|
26
46
|
}
|
|
27
47
|
|
|
48
|
+
/// Is capture wedged rather than merely stumbling?
|
|
49
|
+
///
|
|
50
|
+
/// Deliberately a state and not an event: the daemon reports it, and does
|
|
51
|
+
/// not act on it. A capture loop that restarted the device it is watching
|
|
52
|
+
/// would be a tool that reaches for the mains when a reading looks wrong.
|
|
53
|
+
public var isStalled: Bool {
|
|
54
|
+
reattaches >= Self.stalledAfterReattaches || consecutiveFailures >= Self.stalledAfterFailures
|
|
55
|
+
}
|
|
56
|
+
|
|
28
57
|
public mutating func captureSucceeded() {
|
|
29
58
|
consecutiveFailures = 0
|
|
59
|
+
// A real frame is the only evidence that health is back. Resetting this
|
|
60
|
+
// anywhere else — on a re-resolve, say — is how the loop above stayed
|
|
61
|
+
// invisible.
|
|
62
|
+
reattaches = 0
|
|
30
63
|
}
|
|
31
64
|
|
|
32
65
|
/// Records a failure and says whether the port is now due a re-resolve.
|
|
@@ -50,6 +83,7 @@ public struct CaptureRecovery {
|
|
|
50
83
|
_ = try platform.reattachDisplay()
|
|
51
84
|
try platform.observeChanges(onDamage)
|
|
52
85
|
consecutiveFailures = 0
|
|
86
|
+
reattaches += 1
|
|
53
87
|
return .success(failures)
|
|
54
88
|
} catch {
|
|
55
89
|
// Deliberately not reset: if the port cannot be re-resolved, the
|
|
@@ -80,6 +80,26 @@ public final class FrameStore {
|
|
|
80
80
|
|
|
81
81
|
/// The newest state written, for callers that need the current hashes
|
|
82
82
|
/// without recomputing them.
|
|
83
|
+
/// Publish how capture itself is doing, separately from any frame.
|
|
84
|
+
///
|
|
85
|
+
/// It cannot ride in `state.json`, because that is written when a frame is
|
|
86
|
+
/// recorded and a stall is precisely the absence of frames: the state a
|
|
87
|
+
/// reader sees during a stall is the last healthy one, arbitrarily old, and
|
|
88
|
+
/// there is nothing in it that says so. An idle screen also produces no
|
|
89
|
+
/// frames, so "no frames" is not the signal either — the signal is the
|
|
90
|
+
/// daemon's own failed reads, which only the daemon knows about.
|
|
91
|
+
///
|
|
92
|
+
/// `nil` removes the file: health is the absence of a complaint, so a
|
|
93
|
+
/// reader that finds nothing here is right to assume capture is fine.
|
|
94
|
+
public func writeCaptureHealth(_ health: [String: Any]?) throws {
|
|
95
|
+
let url = root.appendingPathComponent("capture-health.json")
|
|
96
|
+
guard let health else {
|
|
97
|
+
try? FileManager.default.removeItem(at: url)
|
|
98
|
+
return
|
|
99
|
+
}
|
|
100
|
+
try writeAtomic(try JSONSerialization.data(withJSONObject: health), to: url)
|
|
101
|
+
}
|
|
102
|
+
|
|
83
103
|
public func latestState() -> [String: Any]? {
|
|
84
104
|
guard let data = try? Data(contentsOf: root.appendingPathComponent("state.json")) else { return nil }
|
|
85
105
|
return try? JSONSerialization.jsonObject(with: data) as? [String: Any]
|
|
@@ -167,6 +167,7 @@ case "run":
|
|
|
167
167
|
let lock = NSLock()
|
|
168
168
|
var dirty = true
|
|
169
169
|
var recovery = CaptureRecovery()
|
|
170
|
+
var stalledSince: Double?
|
|
170
171
|
var lastCapture = 0.0
|
|
171
172
|
var frames = 0
|
|
172
173
|
var lastReport = Date().timeIntervalSince1970
|
|
@@ -433,7 +434,17 @@ case "run":
|
|
|
433
434
|
latencies.append(Double(DispatchTime.now().uptimeNanoseconds - t0) / 1e6)
|
|
434
435
|
frames += 1
|
|
435
436
|
lastCapture = now
|
|
437
|
+
let wasStalled = recovery.isStalled
|
|
436
438
|
recovery.captureSucceeded()
|
|
439
|
+
if wasStalled {
|
|
440
|
+
// A frame after a stall is the only thing that clears
|
|
441
|
+
// it, and it is worth saying out loud: the device came
|
|
442
|
+
// back on its own, which nobody would otherwise know.
|
|
443
|
+
try? store.writeCaptureHealth(nil)
|
|
444
|
+
stalledSince = nil
|
|
445
|
+
FileHandle.standardError.write(
|
|
446
|
+
"simframed: capture recovered on its own\n".data(using: .utf8)!)
|
|
447
|
+
}
|
|
437
448
|
} catch {
|
|
438
449
|
let due = recovery.captureFailed()
|
|
439
450
|
FileHandle.standardError.write(
|
|
@@ -460,6 +471,25 @@ case "run":
|
|
|
460
471
|
"simframed: could not re-resolve the display port: \(error)\n".data(using: .utf8)!)
|
|
461
472
|
}
|
|
462
473
|
}
|
|
474
|
+
// Say that capture is wedged rather than merely slow, and
|
|
475
|
+
// then do nothing about it. The cure for this state is a
|
|
476
|
+
// device restart, which is the user's to make: a capture
|
|
477
|
+
// loop that rebooted the device it was watching would be a
|
|
478
|
+
// tool reaching for the mains because a reading looked
|
|
479
|
+
// wrong. So it is published, `doctor` grades it and
|
|
480
|
+
// `simframe state` prints it, and an agent reads "the
|
|
481
|
+
// simulator is wedged" instead of "nothing changed".
|
|
482
|
+
if recovery.isStalled {
|
|
483
|
+
if stalledSince == nil { stalledSince = FrameStore.nowMs() }
|
|
484
|
+
try? store.writeCaptureHealth([
|
|
485
|
+
"stalled": true,
|
|
486
|
+
"since": stalledSince ?? FrameStore.nowMs(),
|
|
487
|
+
"at": FrameStore.nowMs(),
|
|
488
|
+
"consecutiveFailures": recovery.consecutiveFailures,
|
|
489
|
+
"reattaches": recovery.reattaches,
|
|
490
|
+
"reason": "\(error)",
|
|
491
|
+
])
|
|
492
|
+
}
|
|
463
493
|
Thread.sleep(forTimeInterval: 0.5)
|
|
464
494
|
}
|
|
465
495
|
}
|
|
@@ -306,6 +306,45 @@ final class CaptureRecoveryTests: XCTestCase {
|
|
|
306
306
|
XCTAssertTrue(damaged, "the damage callback was re-armed on the new descriptor")
|
|
307
307
|
}
|
|
308
308
|
|
|
309
|
+
func testReResolvingTwiceWithoutAFrameIsAStall() {
|
|
310
|
+
// The pathology this is for: re-resolving *works* and reads keep
|
|
311
|
+
// failing. Because a successful re-resolve clears the failure count,
|
|
312
|
+
// the loop is six-failures-then-re-resolve for as long as you let it,
|
|
313
|
+
// and no count of consecutive failures ever notices. Observed three
|
|
314
|
+
// times in one afternoon; only a device restart cured it.
|
|
315
|
+
let platform = StubPlatform()
|
|
316
|
+
_ = try? platform.attach(udid: "STUB-1")
|
|
317
|
+
var recovery = CaptureRecovery(threshold: 2)
|
|
318
|
+
|
|
319
|
+
_ = recovery.captureFailed()
|
|
320
|
+
XCTAssertTrue(recovery.captureFailed())
|
|
321
|
+
_ = recovery.reattach(platform: platform, onDamage: {})
|
|
322
|
+
XCTAssertFalse(recovery.isStalled, "one re-resolve is a recovery, not a stall")
|
|
323
|
+
|
|
324
|
+
_ = recovery.captureFailed()
|
|
325
|
+
XCTAssertTrue(recovery.captureFailed())
|
|
326
|
+
_ = recovery.reattach(platform: platform, onDamage: {})
|
|
327
|
+
XCTAssertTrue(recovery.isStalled, "the second says the port was never the problem")
|
|
328
|
+
|
|
329
|
+
// And only a real frame clears it. Nothing else is evidence.
|
|
330
|
+
recovery.captureSucceeded()
|
|
331
|
+
XCTAssertFalse(recovery.isStalled)
|
|
332
|
+
XCTAssertEqual(recovery.reattaches, 0)
|
|
333
|
+
}
|
|
334
|
+
|
|
335
|
+
func testAReattachThatKeepsFailingIsAlsoAStall() {
|
|
336
|
+
// The other direction: when the re-resolve itself fails the count does
|
|
337
|
+
// keep growing, because nothing resets it.
|
|
338
|
+
let platform = StubPlatform()
|
|
339
|
+
platform.failReattach = true
|
|
340
|
+
var recovery = CaptureRecovery(threshold: 6)
|
|
341
|
+
for _ in 0..<(CaptureRecovery.stalledAfterFailures - 1) { _ = recovery.captureFailed() }
|
|
342
|
+
XCTAssertFalse(recovery.isStalled)
|
|
343
|
+
_ = recovery.captureFailed()
|
|
344
|
+
XCTAssertTrue(recovery.isStalled)
|
|
345
|
+
XCTAssertEqual(recovery.reattaches, 0, "nothing was ever re-resolved")
|
|
346
|
+
}
|
|
347
|
+
|
|
309
348
|
func testAFailedReattachStaysDueRatherThanWaitingForAnotherSix() {
|
|
310
349
|
let platform = StubPlatform()
|
|
311
350
|
platform.failReattach = true
|
package/package.json
CHANGED
|
@@ -1,11 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.7.2",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
|
-
"description": "Always-warm iOS Simulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
5
|
+
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
|
7
7
|
"ios",
|
|
8
8
|
"simulator",
|
|
9
|
+
"android",
|
|
10
|
+
"emulator",
|
|
9
11
|
"mcp",
|
|
10
12
|
"model-context-protocol",
|
|
11
13
|
"claude",
|
|
@@ -0,0 +1,149 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Why do same-screen revisits disagree? Aggregate, not anecdotal.
|
|
3
|
+
//
|
|
4
|
+
// `eval-fingerprint.mjs` says how far apart the two distributions are and names
|
|
5
|
+
// the single worst pair. That was enough while the gap was wide and stopped
|
|
6
|
+
// being enough the day it narrowed: one pair is an anecdote, and the fix has to
|
|
7
|
+
// be aimed at whatever causes most of the divergence.
|
|
8
|
+
//
|
|
9
|
+
// So this reads an eval's `--out` JSON and classifies every divergent token in
|
|
10
|
+
// every same-screen pair. The token grammar is
|
|
11
|
+
// `role:region[:@slot]:w:h["label"]:x:y#count`, which is enough to say what
|
|
12
|
+
// moved:
|
|
13
|
+
//
|
|
14
|
+
// label the same structure under a different chrome label — a dynamic or
|
|
15
|
+
// misread name, the thing chrome labels were most feared to do
|
|
16
|
+
// role the same geometry and place under a different role — inference
|
|
17
|
+
// flipping, e.g. a field that sometimes gets a rectangle
|
|
18
|
+
// bucket the same key, `#1` on one side and `#many` on the other — a
|
|
19
|
+
// sibling count straddling the boundary
|
|
20
|
+
// anchor the same key at a different quantised x/y — the group's topmost
|
|
21
|
+
// member moved by half a grid cell
|
|
22
|
+
// size the same role and place at a different quantised w/h — usually
|
|
23
|
+
// OCR returning a different bounding box for the same text
|
|
24
|
+
// presence a structural key on one side only — an element that came or went
|
|
25
|
+
import fs from 'node:fs';
|
|
26
|
+
import * as fingerprint from '../src/fingerprint.js';
|
|
27
|
+
|
|
28
|
+
const file = process.argv[2];
|
|
29
|
+
if (!file) {
|
|
30
|
+
console.error('usage: node scripts/analyse-fingerprint.mjs <eval --out file.json>');
|
|
31
|
+
process.exit(2);
|
|
32
|
+
}
|
|
33
|
+
const data = JSON.parse(fs.readFileSync(file, 'utf8'));
|
|
34
|
+
const readings = data.readings ?? [];
|
|
35
|
+
if (!readings.length || !readings[0].tokens) {
|
|
36
|
+
console.error('that eval was written without tokens — re-run the eval, it keeps them now');
|
|
37
|
+
process.exit(2);
|
|
38
|
+
}
|
|
39
|
+
|
|
40
|
+
/**
|
|
41
|
+
* A token, taken apart.
|
|
42
|
+
*
|
|
43
|
+
* `role:region[:@slot]:w<n>:h<n>[:"label"]:x<n>:y<n>#<count>` — every field is
|
|
44
|
+
* separated so a pair of tokens can be compared field by field. The first
|
|
45
|
+
* version of this file matched tokens by stripping one field at a time and
|
|
46
|
+
* asking whether the rest was equal, which sounds equivalent and is not: the
|
|
47
|
+
* size-stripped form of a content text token is `text:content`, which matches
|
|
48
|
+
* *every* content text token, so whichever candidate happened to be left over
|
|
49
|
+
* got called a size change. It reported 100% `size` for a sample the eval
|
|
50
|
+
* itself had already shown contained a count flip.
|
|
51
|
+
*/
|
|
52
|
+
const FIELDS = ['role', 'region', 'slot', 'w', 'h', 'label', 'x', 'y', 'count'];
|
|
53
|
+
const TOKEN = /^([^:]+):([^:]+)(?::@([^:]+))?:w(\d+):h(\d+)(?::"([^"]*)")?:x(-?\d+):y(-?\d+)#(1|many)$/;
|
|
54
|
+
|
|
55
|
+
function parse(token) {
|
|
56
|
+
const m = TOKEN.exec(token);
|
|
57
|
+
if (!m) return null;
|
|
58
|
+
const [, role, region, slot, w, h, label, x, y, count] = m;
|
|
59
|
+
return { token, role, region, slot: slot ?? null, w, h, label: label ?? null, x, y, count };
|
|
60
|
+
}
|
|
61
|
+
|
|
62
|
+
/** Which fields differ, in a stable order. */
|
|
63
|
+
function differing(a, b) {
|
|
64
|
+
return FIELDS.filter((f) => a[f] !== b[f]);
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
/**
|
|
68
|
+
* What to call a difference of these fields.
|
|
69
|
+
*
|
|
70
|
+
* Named after the cause rather than the field, because the fix is different for
|
|
71
|
+
* each: a label that moves is normalisation, a role that flips is inference, a
|
|
72
|
+
* count that straddles is bucketing, a box that changes is OCR segmentation.
|
|
73
|
+
*/
|
|
74
|
+
function nameOf(fields) {
|
|
75
|
+
if (!fields.length) return 'identical';
|
|
76
|
+
const names = new Set();
|
|
77
|
+
for (const f of fields) {
|
|
78
|
+
if (f === 'label') names.add('label');
|
|
79
|
+
else if (f === 'role') names.add('role');
|
|
80
|
+
else if (f === 'count') names.add('bucket');
|
|
81
|
+
else if (f === 'x' || f === 'y') names.add('anchor');
|
|
82
|
+
else if (f === 'w' || f === 'h') names.add('size');
|
|
83
|
+
else names.add(f);
|
|
84
|
+
}
|
|
85
|
+
return [...names].sort().join('+');
|
|
86
|
+
}
|
|
87
|
+
|
|
88
|
+
/** Beyond this many differing fields, two tokens are not the same thing moved. */
|
|
89
|
+
const RELATED_MAX_FIELDS = 3;
|
|
90
|
+
|
|
91
|
+
const causes = new Map();
|
|
92
|
+
const byRegion = new Map();
|
|
93
|
+
const bump = (map, k, n = 1) => map.set(k, (map.get(k) ?? 0) + n);
|
|
94
|
+
|
|
95
|
+
const pairs = [];
|
|
96
|
+
for (let i = 0; i < readings.length; i += 1) {
|
|
97
|
+
for (let j = i + 1; j < readings.length; j += 1) {
|
|
98
|
+
if (readings[i].name !== readings[j].name) continue;
|
|
99
|
+
const a = readings[i];
|
|
100
|
+
const b = readings[j];
|
|
101
|
+
const similarity = fingerprint.similarity(a.tokens, b.tokens);
|
|
102
|
+
const setA = new Set(a.tokens);
|
|
103
|
+
const setB = new Set(b.tokens);
|
|
104
|
+
const onlyA = a.tokens.filter((t) => !setB.has(t)).map(parse).filter(Boolean);
|
|
105
|
+
const onlyB = b.tokens.filter((t) => !setA.has(t)).map(parse).filter(Boolean);
|
|
106
|
+
pairs.push({ name: a.name, ra: a.round, rb: b.round, similarity, divergent: onlyA.length + onlyB.length });
|
|
107
|
+
|
|
108
|
+
// Best match, not first match: each unmatched token on the left is paired
|
|
109
|
+
// with the candidate on the right that differs in the fewest fields, and
|
|
110
|
+
// that pairing is consumed. Greedy-by-preference-order is what produced the
|
|
111
|
+
// wrong answer above.
|
|
112
|
+
const remaining = [...onlyB];
|
|
113
|
+
for (const left of onlyA) {
|
|
114
|
+
let best = null;
|
|
115
|
+
for (let k = 0; k < remaining.length; k += 1) {
|
|
116
|
+
const fields = differing(left, remaining[k]);
|
|
117
|
+
if (!best || fields.length < best.fields.length) best = { k, fields };
|
|
118
|
+
}
|
|
119
|
+
const related = best && best.fields.length <= RELATED_MAX_FIELDS;
|
|
120
|
+
const cause = related ? nameOf(best.fields) : 'presence';
|
|
121
|
+
if (related) remaining.splice(best.k, 1);
|
|
122
|
+
bump(causes, cause);
|
|
123
|
+
bump(byRegion, `${left.region}/${left.role}`);
|
|
124
|
+
}
|
|
125
|
+
for (const right of remaining) {
|
|
126
|
+
bump(causes, 'presence');
|
|
127
|
+
bump(byRegion, `${right.region}/${right.role}`);
|
|
128
|
+
}
|
|
129
|
+
}
|
|
130
|
+
}
|
|
131
|
+
|
|
132
|
+
pairs.sort((x, y) => x.similarity - y.similarity);
|
|
133
|
+
const totalDivergent = [...causes.values()].reduce((a, b) => a + b, 0);
|
|
134
|
+
const pct = (n) => `${((n / totalDivergent) * 100).toFixed(0)}%`;
|
|
135
|
+
|
|
136
|
+
console.log(`${pairs.length} same-screen pairs from ${readings.length} readings (${data.label ?? 'unlabelled'}, ${data.device ?? '?'})`);
|
|
137
|
+
console.log(`similarity: min ${pairs[0]?.similarity.toFixed(2)} median ${pairs[pairs.length >> 1]?.similarity.toFixed(2)} max ${pairs[pairs.length - 1]?.similarity.toFixed(2)}`);
|
|
138
|
+
console.log(`\n${totalDivergent} divergent token(s) across those pairs, by cause:`);
|
|
139
|
+
for (const [cause, n] of [...causes].sort((a, b) => b[1] - a[1])) {
|
|
140
|
+
console.log(` ${String(n).padStart(4)} ${pct(n).padStart(4)} ${cause}`);
|
|
141
|
+
}
|
|
142
|
+
console.log('\nby region/role — where the instability lives:');
|
|
143
|
+
for (const [where, n] of [...byRegion].sort((a, b) => b[1] - a[1])) {
|
|
144
|
+
console.log(` ${String(n).padStart(4)} ${pct(n).padStart(4)} ${where}`);
|
|
145
|
+
}
|
|
146
|
+
console.log('\nweakest pairs:');
|
|
147
|
+
for (const p of pairs.slice(0, 8)) {
|
|
148
|
+
console.log(` ${p.similarity.toFixed(2)} ${p.name} r${p.ra} vs r${p.rb} (${p.divergent} divergent)`);
|
|
149
|
+
}
|
package/scripts/bench-flow.mjs
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
import { runScript } from '../src/actions.js';
|
|
4
4
|
import * as input from '../src/input.js';
|
|
5
5
|
import * as api from '../src/index.js';
|
|
6
|
-
import { launchApp, terminateApp } from '../src/
|
|
6
|
+
import { launchApp, terminateApp } from '../src/platform/index.js';
|
|
7
7
|
|
|
8
8
|
const BUNDLE = process.argv[2];
|
|
9
9
|
const device = process.argv[3];
|
|
@@ -123,4 +123,19 @@ if (missing.length) {
|
|
|
123
123
|
);
|
|
124
124
|
process.exit(1);
|
|
125
125
|
}
|
|
126
|
+
// The MCP Registry rejects a description over 100 characters, and it does so at
|
|
127
|
+
// `mcp-publisher validate` — after the tag is pushed and the release has begun.
|
|
128
|
+
// A 0.7.1 release found that out the hard way with a 113-character one. The
|
|
129
|
+
// constraint is the registry's; discovering it locally is this script's job.
|
|
130
|
+
const DESCRIPTION_MAX = 100;
|
|
131
|
+
const server = JSON.parse(fs.readFileSync(path.join(ROOT, 'server.json'), 'utf8'));
|
|
132
|
+
if (server.description.length > DESCRIPTION_MAX) {
|
|
133
|
+
console.error(
|
|
134
|
+
`server.json description is ${server.description.length} characters; the MCP Registry accepts ` +
|
|
135
|
+
`${DESCRIPTION_MAX}. It would fail at validate, with the tag already pushed.`,
|
|
136
|
+
);
|
|
137
|
+
process.exit(1);
|
|
138
|
+
}
|
|
139
|
+
console.log(`server.json description ${server.description.length}/${DESCRIPTION_MAX} chars`);
|
|
140
|
+
|
|
126
141
|
console.log('ok — every build input ships');
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -337,7 +337,19 @@ const novelVerdicts = novelSteps.length
|
|
|
337
337
|
// has not told us anything about the graph.
|
|
338
338
|
const novelRan = novelSteps.length > 0 && novelSteps.every((r) => r.ok !== false);
|
|
339
339
|
check(novelRan, 'the novel action ran at all', `[${novelVerdicts.join(', ')}]`);
|
|
340
|
+
// The other half of the precondition, which was written above as a comment and
|
|
341
|
+
// then trusted. It is not trustworthy: the positioning run sends the device
|
|
342
|
+
// home, and a simulator that has been driven hard stops delivering `home` while
|
|
343
|
+
// still reporting success (docs/DEFERRED.md). From a screen the action cannot
|
|
344
|
+
// change, `no-visible-change` is the honest verdict and the claim below was
|
|
345
|
+
// never asked — so this is a precondition, and saying otherwise is how this
|
|
346
|
+
// file has spent the day accusing the graph of something the device did.
|
|
347
|
+
const novelMoved = novelRan && !novelVerdicts.every((v) => v === 'no-visible-change');
|
|
340
348
|
if (novelRan) {
|
|
349
|
+
check(novelMoved, 'and the device was somewhere the novel action could change',
|
|
350
|
+
novelMoved ? `[${novelVerdicts.join(', ')}]` : 'the screen never moved — the device was already there, or ignored being sent home');
|
|
351
|
+
}
|
|
352
|
+
if (novelRan && novelMoved) {
|
|
341
353
|
check(novelSteps.some((r) => r.verification?.verdict === 'unverified'),
|
|
342
354
|
'an action never taken here before is reported as unverified, not as verified',
|
|
343
355
|
`[${novelVerdicts.join(', ')}]`);
|