simframe 0.12.2 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +123 -9
- package/native/supervise.swift +63 -4
- package/package.json +1 -1
- package/scripts/article-md.mjs +185 -0
- package/scripts/ci-integration-local.sh +22 -1
- package/scripts/ci-memory.mjs +92 -14
- package/scripts/eval-fingerprint.mjs +20 -1
- package/scripts/replay-rulings.mjs +60 -1
- package/src/actions.js +96 -9
- package/src/analyze.js +56 -0
- package/src/cli.js +182 -13
- package/src/fingerprint.js +10 -1
- package/src/index.js +302 -16
- package/src/input.js +4 -0
- package/src/mcp.js +7 -1
- package/src/metrics.js +49 -6
- package/src/ollama.js +37 -9
- package/src/platform/ios.js +30 -5
- package/src/refs.js +12 -1
- package/src/regions.js +54 -0
- package/src/screenmap.js +85 -6
- package/src/store.js +53 -0
- package/src/supervisor.js +20 -2
- package/src/view.js +84 -9
package/README.md
CHANGED
|
@@ -30,6 +30,41 @@ flows run in one call, screens the agent has seen before are answered from
|
|
|
30
30
|
memory, and every answer is text with tap points in it. Nothing returns an
|
|
31
31
|
image unless you ask for one.
|
|
32
32
|
|
|
33
|
+
### It is also a study case, and the numbers are the point
|
|
34
|
+
|
|
35
|
+
Almost everything here was decided by a measurement rather than by an argument,
|
|
36
|
+
and several of those measurements **reversed a decision that had already
|
|
37
|
+
shipped**. They are written down in full, with what we expected beforehand, what
|
|
38
|
+
it cost to find out, and the mistakes made getting there — because a benchmark
|
|
39
|
+
that only records the winner teaches nothing.
|
|
40
|
+
|
|
41
|
+
| | |
|
|
42
|
+
| --- | --- |
|
|
43
|
+
| [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | the record: N, median, p95, machine, every round |
|
|
44
|
+
| [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
|
|
45
|
+
| [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
|
|
46
|
+
| [`docs/ARTICLE.md`](docs/ARTICLE.md) | the argument the whole thing adds up to |
|
|
47
|
+
|
|
48
|
+
The headline results, all free to reuse:
|
|
49
|
+
|
|
50
|
+
- **A warm transition graph removes ~27% of an agent's model round trips** —
|
|
51
|
+
isolated with a three-pass field experiment, and corroborated to three points
|
|
52
|
+
by an independent estimate. Operator knowledge is worth twice that, and no
|
|
53
|
+
tool here carries it yet. ([§14](docs/EXPERIMENTS.md))
|
|
54
|
+
- **A bigger local judge is not a better one.** `qwen3:14b` scored *lower* than
|
|
55
|
+
`qwen3:8b` — 82% against 91% — while being 79% larger and 62% slower. A
|
|
56
|
+
one-line threshold on a number already computed beat all of them at **95% and
|
|
57
|
+
zero latency**.
|
|
58
|
+
- **Every interactive element came from the accessibility tree — 72 of 72 — and
|
|
59
|
+
83% of them have no text at all.** Screen *recognition*, though, did not need
|
|
60
|
+
the tree once. The ladder splits there: semantics need a tree, identity needs
|
|
61
|
+
pixels.
|
|
62
|
+
- **Capture is ~31× faster than a screenshot** for the same work, and the
|
|
63
|
+
capture primitive itself about a thousand times faster.
|
|
64
|
+
- **Three independent reporters converged on one thing**: a refusal is cheap and
|
|
65
|
+
a confident wrong answer is expensive. Every serious bug they found was a
|
|
66
|
+
component reporting more certainty than it had.
|
|
67
|
+
|
|
33
68
|
## What changed, measured
|
|
34
69
|
|
|
35
70
|
Same four-tab navigation flow, on a real production app:
|
|
@@ -177,13 +212,30 @@ was otherwise impossible.
|
|
|
177
212
|
|
|
178
213
|
**What is measured and what is not.** The ranker: 5 of 6 top-1 on hand-written
|
|
179
214
|
cases, median 564 ms warm, and on a real exploration it went to the right region
|
|
180
|
-
in two steps where reading order wandered into version strings.
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
215
|
+
in two steps where reading order wandered into version strings. `ax-first` made
|
|
216
|
+
no measurable difference to how an agent drove a real app, with one small
|
|
217
|
+
regression and one small win.
|
|
218
|
+
|
|
219
|
+
**The supervisor, on 22 labelled failures from a seeded React Native app** —
|
|
220
|
+
and the number to judge it by is round trips, not accuracy:
|
|
221
|
+
|
|
222
|
+
| | |
|
|
223
|
+
|---|---|
|
|
224
|
+
| handled locally, no round trip | **9 of 22** |
|
|
225
|
+
| escalated to Claude | 13 of 22 |
|
|
226
|
+
| a model round trip, measured in the field | 10–16 s |
|
|
227
|
+
|
|
228
|
+
So roughly 90–145 s saved on that population. Accuracy was 77 / 82 / 86% across
|
|
229
|
+
three runs of the *same* questions — it is **not deterministic**, and every
|
|
230
|
+
earlier single-run figure in this project carried that spread without reporting
|
|
231
|
+
it. All of its errors were `wait` where `stop` was right, which is the cheap
|
|
232
|
+
direction: a wrong `wait` costs a settle and a re-run, and `stop` already means
|
|
233
|
+
"hand back to Claude with the unattempted steps" rather than "give up".
|
|
234
|
+
|
|
235
|
+
**Still a bench, not the field.** These are fixtures we designed, on a testbed we
|
|
236
|
+
built, labelled by the person who then scored them. Numbers and conditions are in
|
|
237
|
+
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md); the judgements, including a phase
|
|
238
|
+
cancelled by its own measurement, are in [`docs/DECISIONS.md`](docs/DECISIONS.md).
|
|
187
239
|
|
|
188
240
|
**And the measurements that changed our minds** are in
|
|
189
241
|
[`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md), with what we expected beforehand
|
|
@@ -212,6 +264,15 @@ The Ollama arm is an **experiment, not a recommendation**: off unless named,
|
|
|
212
264
|
no weights shipped, no dependency added, and every arm reads the same briefing
|
|
213
265
|
out of `native/supervise.swift` so no arm is answering a different question.
|
|
214
266
|
|
|
267
|
+
**Two further results, both negative, both worth more than the table.** A
|
|
268
|
+
*cascade* — free rule first, model where it is unsure, Claude after that — was
|
|
269
|
+
worse at every abstention band we tried (95% for the rule alone; 91%, 86%, 82%
|
|
270
|
+
as more was handed to the model). And giving the model a fourth word, `abstain`,
|
|
271
|
+
cost the Apple arm about a third of its accuracy (77/82/86% → 45/50/55%) **while
|
|
272
|
+
it never used the word once**. One added paragraph, nothing else changed. So the
|
|
273
|
+
fourth word ships off, and "add an abstain token" became "find a judge that will
|
|
274
|
+
use one".
|
|
275
|
+
|
|
215
276
|
The supervisor's whole vocabulary is three words on purpose. It cannot invent a
|
|
216
277
|
step, skip one, substitute a target or continue past an unexpected screen — not
|
|
217
278
|
because a threshold forbids it but because those are not answers it can give.
|
|
@@ -538,6 +599,29 @@ simframe flow save checkout ./checkout.json
|
|
|
538
599
|
simframe flow run checkout
|
|
539
600
|
```
|
|
540
601
|
|
|
602
|
+
### What the memory is worth, isolated
|
|
603
|
+
|
|
604
|
+
Three agent sessions drove the same task family on the same production app. The
|
|
605
|
+
third existed only to hold a variable still: pass 2 beat pass 1 by twenty-one
|
|
606
|
+
model round trips, but it had a warm graph *and* an operator who had already
|
|
607
|
+
driven the app once, so the gap could belong to either.
|
|
608
|
+
|
|
609
|
+
| pass | operator | graph | round trips |
|
|
610
|
+
| --- | --- | --- | --- |
|
|
611
|
+
| 1 | fresh | cold | **33** |
|
|
612
|
+
| 3 | fresh | **warm** | **24** |
|
|
613
|
+
| 2 | experienced | warm | **12** |
|
|
614
|
+
|
|
615
|
+
**The graph is worth ~27%** of the round trips. Operator knowledge is worth the
|
|
616
|
+
rest — half of everything left after the graph had taken its share. A second
|
|
617
|
+
reporter, unable to isolate the graph, derived **~23%** from escalation rate per
|
|
618
|
+
step without knowing that number.
|
|
619
|
+
|
|
620
|
+
One app, one operator per pass, and round trips are not seconds. What it settles
|
|
621
|
+
is the direction: the memory is real, it is the smaller half, and the larger half
|
|
622
|
+
is what an operator learns and no tool here carries between sessions yet. Full
|
|
623
|
+
working in [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) §14.
|
|
624
|
+
|
|
541
625
|
## Does this work on *your* app?
|
|
542
626
|
|
|
543
627
|
Nothing in simframe is written for a particular app. What varies between apps is
|
|
@@ -553,9 +637,27 @@ Run `simframe ui` on any screen to see exactly what simframe can see, with each
|
|
|
553
637
|
target marked `ax` or `ocr`. If something you can read is not listed, that is a
|
|
554
638
|
bug worth reporting.
|
|
555
639
|
|
|
640
|
+
A control the app **declares but never names** — an icon-only overflow menu, a
|
|
641
|
+
back chevron — is listed with its coordinates and tappable by `#ref` rather than
|
|
642
|
+
by name, and the map counts how many such controls are on the screen. A view the
|
|
643
|
+
app never declared accessible at all is invisible to any accessibility tree, ours
|
|
644
|
+
included, and the same line says so: a count of zero on a screen that has one
|
|
645
|
+
would be the more expensive answer.
|
|
646
|
+
|
|
647
|
+
**What a settle can and cannot see.** Stillness is decided from a mean over a
|
|
648
|
+
grid of the frame, and a small animation does not move a mean. Measured on a
|
|
649
|
+
screen built to never settle — a spinner, frames every 77–95 ms — the mean
|
|
650
|
+
difference was **0.00196** against a 0.004 threshold while one cell moved by
|
|
651
|
+
**0.0275**, and `stableForMs` reported **79 seconds** of stillness. So a settle
|
|
652
|
+
can return satisfied while part of the screen is still moving, and it now says
|
|
653
|
+
so: `settled after 63ms (a 13x13 region is still animating)`. A settle that
|
|
654
|
+
times out names where the movement is instead. Refusing to settle on any
|
|
655
|
+
animation would be right for a spinner and wrong for a blinking cursor, so the
|
|
656
|
+
disagreement is reported rather than resolved by a guessed threshold.
|
|
657
|
+
|
|
556
658
|
Two honest caveats. OCR reads **text**, so a purely graphical icon with no label
|
|
557
|
-
is invisible to both paths —
|
|
558
|
-
|
|
659
|
+
is invisible to both paths — the tree still gives you its coordinates, and
|
|
660
|
+
`#ref` still taps it. And the confirm-button vocabulary (`APPLY`, `OK`, `SAVE`,
|
|
559
661
|
`DONE`…) is English; a localised UI needs those words extended.
|
|
560
662
|
|
|
561
663
|
## Measured
|
|
@@ -650,6 +752,18 @@ accepted as the name that loop used to have.
|
|
|
650
752
|
a support banner was swallowing — a banner that was *in the element list the
|
|
651
753
|
same call printed*. Nothing said "there is something at y≈753 and you started
|
|
652
754
|
at y=750".
|
|
755
|
+
- **A control half off the edge is offered, clamped, and says so.** A filter
|
|
756
|
+
chip with 29 pt of itself visible used to be dropped — the off-screen test
|
|
757
|
+
reads the *centre* — and OCR's reading of that same sliver, a box labelled
|
|
758
|
+
`Flc`, was printed in its place. So the map did not merely omit a control; it
|
|
759
|
+
offered a meaningless name for it at an ordinary-looking coordinate. Now:
|
|
760
|
+
|
|
761
|
+
```
|
|
762
|
+
#6 button 387,191 Flowering (partly off-screen — the coordinate is the middle of the visible part)
|
|
763
|
+
```
|
|
764
|
+
|
|
765
|
+
`scrollTo` is deliberately unchanged: a clipped element is still "not in
|
|
766
|
+
view", so it still scrolls rather than calling a sliver good enough.
|
|
653
767
|
|
|
654
768
|
## CLI
|
|
655
769
|
|
package/native/supervise.swift
CHANGED
|
@@ -40,7 +40,40 @@ struct Judgement {
|
|
|
40
40
|
var decision: Decision
|
|
41
41
|
}
|
|
42
42
|
|
|
43
|
+
/**
|
|
44
|
+
* The same question with a fourth answer: "I cannot tell".
|
|
45
|
+
*
|
|
46
|
+
* A second type rather than a fourth case on `Decision`, because the two have
|
|
47
|
+
* to be *comparable*. `@Generable` fixes the answer space at compile time, so a
|
|
48
|
+
* binary that can only be asked one way cannot measure whether the token helps
|
|
49
|
+
* or merely moves the errors around — and "does the model abstain on the ones
|
|
50
|
+
* it gets wrong, or at random" is the whole question. One binary, both
|
|
51
|
+
* vocabularies, chosen per request.
|
|
52
|
+
*
|
|
53
|
+
* `abstain` is not a new capability and cannot become one. It means "behave as
|
|
54
|
+
* if there is no supervisor", which is exactly the `null` the caller already
|
|
55
|
+
* handles on every failure path. The answer space stays the safety property:
|
|
56
|
+
* this strictly shrinks what the model can cause to happen.
|
|
57
|
+
*/
|
|
58
|
+
@available(macOS 26.0, *)
|
|
59
|
+
@Generable
|
|
60
|
+
enum CautiousDecision: String {
|
|
61
|
+
case wait
|
|
62
|
+
case retry
|
|
63
|
+
case stop
|
|
64
|
+
case abstain
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
@available(macOS 26.0, *)
|
|
68
|
+
@Generable
|
|
69
|
+
struct CautiousJudgement {
|
|
70
|
+
@Guide(description: "wait if the screen is still arriving, retry if the same step should be attempted again, stop if nothing further can work, abstain if you cannot tell from what you were given")
|
|
71
|
+
var decision: CautiousDecision
|
|
72
|
+
}
|
|
73
|
+
|
|
43
74
|
struct Situation: Decodable {
|
|
75
|
+
/// Ask with the fourth word available. Absent means the shipped three.
|
|
76
|
+
let mayAbstain: Bool?
|
|
44
77
|
let goal: String?
|
|
45
78
|
let step: String
|
|
46
79
|
let expected: String?
|
|
@@ -121,6 +154,19 @@ func serve() async {
|
|
|
121
154
|
are not being asked what to do, only whether this can proceed. Answer with \
|
|
122
155
|
the decision alone.
|
|
123
156
|
"""
|
|
157
|
+
|
|
158
|
+
// The same brief with the fourth word, built from the first so the two
|
|
159
|
+
// cannot drift. Everything above applies; this adds one rule and changes
|
|
160
|
+
// nothing else, which is what makes the comparison mean something.
|
|
161
|
+
let abstainInstructions = instructions + "\n\n" + """
|
|
162
|
+
There is a fourth answer: abstain. Use it when the evidence you were \
|
|
163
|
+
given does not distinguish these cases — when you would be guessing. \
|
|
164
|
+
Abstaining is not a failure and it is not penalised: it hands the decision \
|
|
165
|
+
to someone with more context, which costs one step. A wrong stop abandons a \
|
|
166
|
+
plan that would have worked, and a wrong wait spends a timeout on a screen \
|
|
167
|
+
that will never change. Prefer abstain to either. Do not abstain merely \
|
|
168
|
+
because the situation is unusual; abstain because the evidence is absent.
|
|
169
|
+
"""
|
|
124
170
|
// Warm up front so the first real judgement does not pay model load —
|
|
125
171
|
// measured at ~880ms against ~600ms warm.
|
|
126
172
|
//
|
|
@@ -133,7 +179,17 @@ func serve() async {
|
|
|
133
179
|
// The window, reported rather than assumed. 4,096 has been a documented
|
|
134
180
|
// constant we repeated; since 26.4 it is queryable, so it is now read from
|
|
135
181
|
// the model and handed to the caller, who prints it in `doctor`.
|
|
136
|
-
|
|
182
|
+
// Guarded at 26.4, which is where `contextSize` became queryable — the
|
|
183
|
+
// enclosing block only promises 26.0, so an SDK between the two would fail
|
|
184
|
+
// to *compile* and the supervisor would simply not exist on that machine,
|
|
185
|
+
// reported as "could not build the local supervisor". A peer round is
|
|
186
|
+
// exactly where that costs somebody an afternoon, and the window is a
|
|
187
|
+
// nicety: it is printed in `doctor` and nothing depends on it.
|
|
188
|
+
var hello: [String: Any] = ["ready": true]
|
|
189
|
+
if #available(macOS 26.4, *) {
|
|
190
|
+
hello["contextSize"] = SystemLanguageModel.default.contextSize
|
|
191
|
+
}
|
|
192
|
+
emit(hello)
|
|
137
193
|
while let line = readLine(strippingNewline: true) {
|
|
138
194
|
if line.isEmpty { continue }
|
|
139
195
|
guard let data = line.data(using: .utf8),
|
|
@@ -157,8 +213,11 @@ func serve() async {
|
|
|
157
213
|
}
|
|
158
214
|
let started = Date()
|
|
159
215
|
do {
|
|
160
|
-
let
|
|
161
|
-
let
|
|
216
|
+
let mayAbstain = s.mayAbstain ?? false
|
|
217
|
+
let session = LanguageModelSession(instructions: mayAbstain ? abstainInstructions : instructions)
|
|
218
|
+
let decision: String = mayAbstain
|
|
219
|
+
? try await session.respond(to: prompt, generating: CautiousJudgement.self).content.decision.rawValue
|
|
220
|
+
: try await session.respond(to: prompt, generating: Judgement.self).content.decision.rawValue
|
|
162
221
|
// No reason field, deliberately. Asked for one it confabulated in
|
|
163
222
|
// every observed run: a correct `stop` justified as "screen is
|
|
164
223
|
// elsewhere" when the screen was exactly where the plan expected,
|
|
@@ -170,7 +229,7 @@ func serve() async {
|
|
|
170
229
|
// one."* What the caller gets instead is which rule or which model
|
|
171
230
|
// answered, which is true by construction.
|
|
172
231
|
emit([
|
|
173
|
-
"decision":
|
|
232
|
+
"decision": decision,
|
|
174
233
|
"ms": Int(Date().timeIntervalSince(started) * 1000),
|
|
175
234
|
])
|
|
176
235
|
} catch let err as LanguageModelSession.GenerationError {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.13.0",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
5
|
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
|
@@ -0,0 +1,185 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Generate docs/ARTICLE.md from docs/agents-shouldnt-blink.html.
|
|
3
|
+
//
|
|
4
|
+
// node scripts/article-md.mjs # write it
|
|
5
|
+
// node scripts/article-md.mjs --check # fail if it is out of date
|
|
6
|
+
//
|
|
7
|
+
// **Why generated rather than kept in step by hand.** The two were written as a
|
|
8
|
+
// pair, with a note at the top of the Markdown saying it was the source of
|
|
9
|
+
// record and to edit both together. Then the page grew several sections the
|
|
10
|
+
// Markdown never got — the local supervisor, the capacity comparison — and the
|
|
11
|
+
// note stayed there being wrong, which is worse than no note: the next person
|
|
12
|
+
// follows it. Replacing a stale instruction with a stale *warning* would have
|
|
13
|
+
// been the same mistake one level up.
|
|
14
|
+
//
|
|
15
|
+
// So the page is the source and this is the projection. Two copies of a
|
|
16
|
+
// document cannot drift when one of them is derived, and `--check` in CI means
|
|
17
|
+
// they cannot drift silently even for a commit.
|
|
18
|
+
//
|
|
19
|
+
// This is not a general HTML-to-Markdown converter and should not become one.
|
|
20
|
+
// It understands exactly the constructs that article uses, and it *throws* on
|
|
21
|
+
// anything it does not recognise rather than dropping it — a converter that
|
|
22
|
+
// silently skips an unfamiliar tag produces a document that looks complete and
|
|
23
|
+
// is not, which is the failure mode this whole file exists to prevent.
|
|
24
|
+
import fs from 'node:fs';
|
|
25
|
+
import path from 'node:path';
|
|
26
|
+
import { fileURLToPath } from 'node:url';
|
|
27
|
+
|
|
28
|
+
const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
|
|
29
|
+
const SOURCE = path.join(ROOT, 'docs', 'agents-shouldnt-blink.html');
|
|
30
|
+
const TARGET = path.join(ROOT, 'docs', 'ARTICLE.md');
|
|
31
|
+
|
|
32
|
+
const ENTITIES = {
|
|
33
|
+
mdash: '—', ndash: '–', nbsp: ' ', amp: '&', lt: '<', gt: '>', quot: '"',
|
|
34
|
+
rsquo: '’', lsquo: '‘', ldquo: '“', rdquo: '”',
|
|
35
|
+
times: '×', rarr: '→', larr: '←', hellip: '…', plusmn: '±', deg: '°',
|
|
36
|
+
middot: '·', bull: '•', frac12: '½', le: '≤', ge: '≥', minus: '−', apos: "'",
|
|
37
|
+
};
|
|
38
|
+
|
|
39
|
+
/** Entities, then inline tags, then whitespace. Order matters. */
|
|
40
|
+
function inline(html) {
|
|
41
|
+
let out = html
|
|
42
|
+
// The faculty spans open a paragraph and name the thing it defines —
|
|
43
|
+
// "Eyes", "Hand", "Memory" — with no separator in the HTML because the
|
|
44
|
+
// stylesheet supplies one.
|
|
45
|
+
.replace(/<span class="faculty">([\s\S]*?)<\/span>/g, (_, t) => `<strong>${t}</strong> — `)
|
|
46
|
+
// An abbreviation carries its expansion in the title attribute, which a
|
|
47
|
+
// reader of the Markdown has no way to hover over. Keep both.
|
|
48
|
+
.replace(/<abbr title="([^"]*)"[^>]*>([\s\S]*?)<\/abbr>/g, (_, full, t) => `${stripTags(t)} (${full})`)
|
|
49
|
+
.replace(/<a\b[^>]*href="([^"]*)"[^>]*>([\s\S]*?)<\/a>/g, (_, href, text) => `[${inline(text)}](${href})`)
|
|
50
|
+
.replace(/<code>([\s\S]*?)<\/code>/g, (_, t) => `\`${stripTags(t)}\``)
|
|
51
|
+
.replace(/<(?:b|strong)>([\s\S]*?)<\/(?:b|strong)>/g, (_, t) => `**${inline(t)}**`)
|
|
52
|
+
.replace(/<(?:i|em)>([\s\S]*?)<\/(?:i|em)>/g, (_, t) => `*${inline(t)}*`)
|
|
53
|
+
.replace(/<br\s*\/?>/g, ' \n');
|
|
54
|
+
const left = out.match(/<(?!\/)[a-z]/i);
|
|
55
|
+
if (left) throw new Error(`unhandled inline tag near: ${out.slice(Math.max(0, left.index - 40), left.index + 60)}`);
|
|
56
|
+
return decode(out).replace(/\s+/g, ' ').trim();
|
|
57
|
+
}
|
|
58
|
+
|
|
59
|
+
const decode = (s) => s
|
|
60
|
+
.replace(/&#(\d+);/g, (_, n) => String.fromCodePoint(Number(n)))
|
|
61
|
+
.replace(/&([a-zA-Z0-9]+);/g, (m, name) => {
|
|
62
|
+
if (!(name in ENTITIES)) throw new Error(`unknown entity &${name};`);
|
|
63
|
+
return ENTITIES[name];
|
|
64
|
+
});
|
|
65
|
+
|
|
66
|
+
const stripTags = (s) => decode(s.replace(/<[^>]+>/g, '')).replace(/\s+/g, ' ').trim();
|
|
67
|
+
|
|
68
|
+
/** A table, as GitHub-flavoured Markdown. */
|
|
69
|
+
function table(html) {
|
|
70
|
+
const rows = [...html.matchAll(/<tr>([\s\S]*?)<\/tr>/g)].map((m) =>
|
|
71
|
+
[...m[1].matchAll(/<(th|td)\b[^>]*>([\s\S]*?)<\/\1>/g)].map((c) => inline(c[2])));
|
|
72
|
+
if (!rows.length) return [];
|
|
73
|
+
const head = /<thead>/.test(html) ? rows[0] : rows[0].map(() => ' ');
|
|
74
|
+
const body = /<thead>/.test(html) ? rows.slice(1) : rows;
|
|
75
|
+
const width = Math.max(...rows.map((r) => r.length));
|
|
76
|
+
const pad = (r) => [...r, ...Array(width - r.length).fill('')];
|
|
77
|
+
return [
|
|
78
|
+
`| ${pad(head).join(' | ')} |`,
|
|
79
|
+
`| ${Array(width).fill('---').join(' | ')} |`,
|
|
80
|
+
...body.map((r) => `| ${pad(r).join(' | ')} |`),
|
|
81
|
+
];
|
|
82
|
+
}
|
|
83
|
+
|
|
84
|
+
/** Everything inside one `<section>`, in order. */
|
|
85
|
+
function blocks(html) {
|
|
86
|
+
const out = [];
|
|
87
|
+
// One pass over the constructs the article actually uses. `lastIndex`
|
|
88
|
+
// walking rather than nested parsing, because the document is flat inside a
|
|
89
|
+
// section and a real parser here would be a dependency.
|
|
90
|
+
const re = new RegExp([
|
|
91
|
+
'<h([123])>([\\s\\S]*?)<\\/h\\1>',
|
|
92
|
+
'<p class="sec-label">([\\s\\S]*?)<\\/p>',
|
|
93
|
+
'<p class="panel-label">([\\s\\S]*?)<\\/p>',
|
|
94
|
+
'<span class="amend-label">([\\s\\S]*?)<\\/span>',
|
|
95
|
+
'<p\\b[^>]*>([\\s\\S]*?)<\\/p>',
|
|
96
|
+
'<table>([\\s\\S]*?)<\\/table>',
|
|
97
|
+
'<pre>([\\s\\S]*?)<\\/pre>',
|
|
98
|
+
'<li>([\\s\\S]*?)<\\/li>',
|
|
99
|
+
].join('|'), 'g');
|
|
100
|
+
let m;
|
|
101
|
+
while ((m = re.exec(html)) !== null) {
|
|
102
|
+
const [, level, heading, sec, panel, amend, para, tbl, pre, li] = m;
|
|
103
|
+
if (heading != null) out.push({ kind: 'h', level: Number(level), text: inline(heading) });
|
|
104
|
+
else if (sec != null) out.push({ kind: 'kicker', text: inline(sec) });
|
|
105
|
+
else if (panel != null) out.push({ kind: 'label', text: inline(panel) });
|
|
106
|
+
else if (amend != null) out.push({ kind: 'label', text: inline(amend) });
|
|
107
|
+
else if (para != null) out.push({ kind: 'p', text: inline(para) });
|
|
108
|
+
else if (tbl != null) out.push({ kind: 'table', lines: table(tbl) });
|
|
109
|
+
else if (pre != null) out.push({ kind: 'pre', text: decode(stripTagsPreserving(pre)) });
|
|
110
|
+
else if (li != null) out.push({ kind: 'li', text: inline(li) });
|
|
111
|
+
}
|
|
112
|
+
return out;
|
|
113
|
+
}
|
|
114
|
+
|
|
115
|
+
/** `<pre>` keeps its newlines; only tags come out. */
|
|
116
|
+
const stripTagsPreserving = (s) => s.replace(/<[^>]+>/g, '').replace(/^\n/, '').replace(/\s+$/, '');
|
|
117
|
+
|
|
118
|
+
export function render(html) {
|
|
119
|
+
const title = stripTags(html.match(/<h1>([\s\S]*?)<\/h1>/)[1]);
|
|
120
|
+
const kicker = stripTags(html.match(/<p class="kicker">([\s\S]*?)<\/p>/)[1]);
|
|
121
|
+
const standfirst = html.match(/<p class="standfirst">([\s\S]*?)<\/p>/);
|
|
122
|
+
|
|
123
|
+
const lines = [
|
|
124
|
+
`# ${title}`,
|
|
125
|
+
'',
|
|
126
|
+
`*${kicker}*`,
|
|
127
|
+
'',
|
|
128
|
+
'> **This file is generated** from [`agents-shouldnt-blink.html`](agents-shouldnt-blink.html)',
|
|
129
|
+
'> by `scripts/article-md.mjs`. Edit the page, not this — a hand-kept copy of a',
|
|
130
|
+
'> document is a copy that drifts, and this one did.',
|
|
131
|
+
'>',
|
|
132
|
+
'> The measurements, with N, median, p95 and the mistakes made getting to each,',
|
|
133
|
+
'> are in [`BENCHMARKS.md`](BENCHMARKS.md); what we expected before measuring is',
|
|
134
|
+
'> in [`EXPERIMENTS.md`](EXPERIMENTS.md).',
|
|
135
|
+
'',
|
|
136
|
+
];
|
|
137
|
+
if (standfirst) lines.push(inline(standfirst[1]), '');
|
|
138
|
+
lines.push('---', '');
|
|
139
|
+
|
|
140
|
+
for (const section of html.match(/<section>[\s\S]*?<\/section>/g) ?? []) {
|
|
141
|
+
let pendingKicker = null;
|
|
142
|
+
let inList = false;
|
|
143
|
+
for (const b of blocks(section)) {
|
|
144
|
+
if (b.kind !== 'li' && inList) { lines.push(''); inList = false; }
|
|
145
|
+
switch (b.kind) {
|
|
146
|
+
// The kicker precedes its heading on the page and reads as a label for
|
|
147
|
+
// it, so it is held until the heading arrives rather than emitted where
|
|
148
|
+
// it was found.
|
|
149
|
+
case 'kicker': pendingKicker = b.text; break;
|
|
150
|
+
case 'h':
|
|
151
|
+
if (pendingKicker) { lines.push(`*${pendingKicker}*`, ''); pendingKicker = null; }
|
|
152
|
+
// `<h1>` is the title, already emitted above, so the page's `<h2>`
|
|
153
|
+
// sections are the document's second level and not its third.
|
|
154
|
+
lines.push(`${'#'.repeat(b.level)} ${b.text}`, '');
|
|
155
|
+
break;
|
|
156
|
+
case 'label': lines.push(`**${b.text}**`, ''); break;
|
|
157
|
+
case 'p': lines.push(b.text, ''); break;
|
|
158
|
+
case 'li': lines.push(`- ${b.text}`); inList = true; break;
|
|
159
|
+
case 'table': lines.push(...b.lines, ''); break;
|
|
160
|
+
case 'pre': lines.push('```', b.text, '```', ''); break;
|
|
161
|
+
default: throw new Error(`unhandled block ${b.kind}`);
|
|
162
|
+
}
|
|
163
|
+
}
|
|
164
|
+
if (inList) lines.push('');
|
|
165
|
+
}
|
|
166
|
+
|
|
167
|
+
return `${lines.join('\n').replace(/\n{3,}/g, '\n\n').trim()}\n`;
|
|
168
|
+
}
|
|
169
|
+
|
|
170
|
+
const html = fs.readFileSync(SOURCE, 'utf8');
|
|
171
|
+
const md = render(html);
|
|
172
|
+
|
|
173
|
+
if (process.argv.includes('--check')) {
|
|
174
|
+
const current = fs.existsSync(TARGET) ? fs.readFileSync(TARGET, 'utf8') : '';
|
|
175
|
+
if (current === md) {
|
|
176
|
+
console.log(`docs/ARTICLE.md is in step with the page (${md.split('\n').length} lines)`);
|
|
177
|
+
process.exit(0);
|
|
178
|
+
}
|
|
179
|
+
console.error('docs/ARTICLE.md is out of date with docs/agents-shouldnt-blink.html.');
|
|
180
|
+
console.error('Run: node scripts/article-md.mjs');
|
|
181
|
+
process.exit(1);
|
|
182
|
+
}
|
|
183
|
+
|
|
184
|
+
fs.writeFileSync(TARGET, md);
|
|
185
|
+
console.log(`wrote docs/ARTICLE.md — ${md.split('\n').length} lines from ${SOURCE.split('/').pop()}`);
|
|
@@ -43,6 +43,15 @@ for (const key of ["input.driver", "ax.driver"]) {
|
|
|
43
43
|
console.log(`${good ? "ok " : "FAIL"} ${key} = ${JSON.stringify(d[key])} (want "simframed")`);
|
|
44
44
|
if (!good) failed = true;
|
|
45
45
|
}
|
|
46
|
+
// A configured driver is not an answering driver — the same check as ci.yml.
|
|
47
|
+
// A whole CI run read the screen eighteen times, every reading came back
|
|
48
|
+
// OCR-only, and this step said `ok` because a driver was present.
|
|
49
|
+
{
|
|
50
|
+
const n = d["ax.elements"];
|
|
51
|
+
const good = typeof n === "number" && n > 0;
|
|
52
|
+
console.log(`${good ? "ok " : "FAIL"} ax.elements = ${JSON.stringify(n)} (want > 0 — the tree must answer, not merely exist)`);
|
|
53
|
+
if (!good) failed = true;
|
|
54
|
+
}
|
|
46
55
|
if (d.warnings > 0) {
|
|
47
56
|
console.log(`\n${d.warnings} degraded layer(s):`);
|
|
48
57
|
for (const c of d.checks.filter((c) => c.level !== "ok")) console.log(` ${c.level} ${c.name}: ${c.detail}`);
|
|
@@ -68,7 +77,19 @@ else
|
|
|
68
77
|
fi
|
|
69
78
|
|
|
70
79
|
step "The memory layer — screen map, refs, graph, verdicts, flows"
|
|
71
|
-
|
|
80
|
+
# Mirrors the job: exit 75 means the display wedged and nothing was tested, so
|
|
81
|
+
# revive once and run again; any other non-zero is a real check failing. The
|
|
82
|
+
# wedge lands on an app switch and hit roughly every other run of this script on
|
|
83
|
+
# the day it was written, which is most of why the job looked flaky.
|
|
84
|
+
if node scripts/ci-memory.mjs --device="$DEVICE"; then
|
|
85
|
+
ok "ci-memory"
|
|
86
|
+
elif [ $? = 75 ]; then
|
|
87
|
+
printf ' (the display wedged — DEFERRED 126. Reviving once and running again.)\n'
|
|
88
|
+
node src/cli.js revive --device="$DEVICE" || true
|
|
89
|
+
node scripts/ci-memory.mjs --device="$DEVICE" && ok "ci-memory (after one revive)" || bad "ci-memory"
|
|
90
|
+
else
|
|
91
|
+
bad "ci-memory"
|
|
92
|
+
fi
|
|
72
93
|
|
|
73
94
|
step "The fingerprint distributions, measured and bounded"
|
|
74
95
|
node scripts/eval-fingerprint.mjs --tour=test/tours/device-native.json --rounds=3 \
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -50,6 +50,9 @@ let failures = 0;
|
|
|
50
50
|
*/
|
|
51
51
|
let deviceDied = null;
|
|
52
52
|
|
|
53
|
+
/** EX_TEMPFAIL: the device died under the checks, so nothing was tested. */
|
|
54
|
+
const DEVICE_DIED_EXIT = 75;
|
|
55
|
+
|
|
53
56
|
function check(ok, label, detail = '') {
|
|
54
57
|
if (!ok) failures += 1;
|
|
55
58
|
console.log(`${ok ? 'ok ' : 'FAIL'} ${label}${detail ? ` — ${detail}` : ''}`);
|
|
@@ -171,9 +174,21 @@ async function jsonRetry(args, opts, attempts = 3) {
|
|
|
171
174
|
console.error(` ${String(last.message).split('\n')[0]}`);
|
|
172
175
|
console.error('\nEverything after this point would be testing a dead simulator, so the run');
|
|
173
176
|
console.error('stops here. This is not a memory-layer failure — it is the device-state');
|
|
174
|
-
console.error('problem in docs/DEFERRED.md. A device restart is the only known cure
|
|
175
|
-
console.error('
|
|
176
|
-
|
|
177
|
+
console.error('problem in docs/DEFERRED.md (126). A device restart is the only known cure,');
|
|
178
|
+
console.error('and `simframe revive` is that restart.');
|
|
179
|
+
// Exit 75, not 1, and the distinction is the whole point of this file.
|
|
180
|
+
//
|
|
181
|
+
// "A check about the memory layer failed" and "the simulator died under the
|
|
182
|
+
// checks" are different conditions with different responses, and for two CI
|
|
183
|
+
// rounds they were one exit code — so a caller could only retry everything
|
|
184
|
+
// or retry nothing. 75 is EX_TEMPFAIL, which is exactly what this is: the
|
|
185
|
+
// subject under test was never reached.
|
|
186
|
+
//
|
|
187
|
+
// The caller reviving and running again is not papering over a product bug.
|
|
188
|
+
// The wedge is a documented CoreSimulator condition, `frame --fresh` names
|
|
189
|
+
// it, and `revive` is the cure this project ships for it — CI simply had no
|
|
190
|
+
// way to say "use it".
|
|
191
|
+
process.exit(DEVICE_DIED_EXIT);
|
|
177
192
|
}
|
|
178
193
|
throw last;
|
|
179
194
|
}
|
|
@@ -244,7 +259,36 @@ if (failures) {
|
|
|
244
259
|
|
|
245
260
|
console.log('\n--- the screen map ---');
|
|
246
261
|
await jsonRetry(['do', LOOP], { allowFail: true });
|
|
247
|
-
const map = await
|
|
262
|
+
const map = await readableMap();
|
|
263
|
+
|
|
264
|
+
/**
|
|
265
|
+
* A map that is empty *and* says a sensor failed is a blink, not a result.
|
|
266
|
+
*
|
|
267
|
+
* `jsonRetry` retries a command that throws, and this one did not throw: `ui`
|
|
268
|
+
* returned 200 with zero elements and a degraded note, which is the shape of
|
|
269
|
+
* failure this whole repo keeps writing rules about. The run that made this
|
|
270
|
+
* necessary had the accessibility read time out once, report an empty map, and
|
|
271
|
+
* then resolve a ref correctly sixty seconds later on the same device — so the
|
|
272
|
+
* device was fine and the check had caught one bad read.
|
|
273
|
+
*
|
|
274
|
+
* Deliberately narrow. An empty map with **no** degraded sensor is a real
|
|
275
|
+
* answer — that is a blank screen and the check should fail on it. Only an
|
|
276
|
+
* empty map that admits a layer did not answer is worth asking again, and after
|
|
277
|
+
* three attempts it fails with what the sensor said, which is the diagnosis
|
|
278
|
+
* either way.
|
|
279
|
+
*/
|
|
280
|
+
async function readableMap(attempts = 3) {
|
|
281
|
+
let last;
|
|
282
|
+
for (let i = 0; i < attempts; i += 1) {
|
|
283
|
+
last = await jsonRetry(['ui']);
|
|
284
|
+
if (last.elements?.length || !(last.degraded ?? []).length) return last;
|
|
285
|
+
if (i < attempts - 1) {
|
|
286
|
+
console.log(` (the map came back empty and a sensor said why — retrying \`ui\`: ${(last.degraded ?? []).join('; ')})`);
|
|
287
|
+
await new Promise((r) => setTimeout(r, 2000));
|
|
288
|
+
}
|
|
289
|
+
}
|
|
290
|
+
return last;
|
|
291
|
+
}
|
|
248
292
|
|
|
249
293
|
// Report what actually answered rather than asserting the runner's situation.
|
|
250
294
|
// This line used to read "with no accessibility tree available" unconditionally,
|
|
@@ -263,14 +307,38 @@ check(Number.isFinite(map.points?.width) && Number.isFinite(map.points?.height),
|
|
|
263
307
|
'the map knows the screen size in points', `${map.points?.width}x${map.points?.height}pt`);
|
|
264
308
|
|
|
265
309
|
const refs = (map.elements ?? []).map((e) => e.ref);
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
310
|
+
// Three checks over a collection, and `every`/`!some` are all true of an empty
|
|
311
|
+
// one. On a CI run whose map came back with **0 elements** they printed
|
|
312
|
+
// `ok refs are numbered 1..n with no gaps — #1..#0` and two more like it:
|
|
313
|
+
// three lines of reassurance about nothing, directly under the failure that
|
|
314
|
+
// said the map was empty.
|
|
315
|
+
//
|
|
316
|
+
// That is the exact defect three field reports spent a day describing — a
|
|
317
|
+
// confident statement that verified nothing — and the harness was doing it to
|
|
318
|
+
// itself in the same output. Untested is not passed.
|
|
319
|
+
if (!map.elements?.length) {
|
|
320
|
+
for (const label of [
|
|
321
|
+
'refs are numbered 1..n with no gaps',
|
|
322
|
+
'every element has a tap point on the screen',
|
|
323
|
+
'the status bar is not offered as something to tap',
|
|
324
|
+
]) skip(label, 'the map was empty, so there was nothing to check');
|
|
325
|
+
// And say what the device was showing, because an empty map on a live device
|
|
326
|
+
// is the shape this project has chased under four different symptoms. The
|
|
327
|
+
// liveness note carries the ignored-gesture check; `stableForMs` and the
|
|
328
|
+
// frame age are the two numbers whose disagreement names a dead surface.
|
|
329
|
+
const st = await jsonRetry(['state'], { allowFail: true });
|
|
330
|
+
console.log(` the device at that moment: frame #${st?.seq ?? '?'}, `
|
|
331
|
+
+ `${st?.stableForMs ?? '?'}ms still, ${st?.live?.note ?? 'liveness reported nothing'}`);
|
|
332
|
+
} else {
|
|
333
|
+
check(refs.every((r, i) => r === i + 1),
|
|
334
|
+
'refs are numbered 1..n with no gaps', `#1..#${refs.length}`);
|
|
335
|
+
check(map.elements.every((e) =>
|
|
336
|
+
Number.isInteger(e.x) && Number.isInteger(e.y)
|
|
337
|
+
&& e.y >= 0 && e.y <= map.points.height && e.x >= 0 && e.x <= map.points.width),
|
|
338
|
+
'every element has a tap point on the screen');
|
|
339
|
+
check(!map.elements.some((e) => e.region === 'status-bar'),
|
|
340
|
+
'the status bar is not offered as something to tap');
|
|
341
|
+
}
|
|
274
342
|
|
|
275
343
|
console.log('\n--- element refs ---');
|
|
276
344
|
// Re-read on a screen we chose, rather than on whatever the device happened to
|
|
@@ -296,7 +364,8 @@ if (first) {
|
|
|
296
364
|
// hashes only have to agree with that, and they only get a say when they are
|
|
297
365
|
// informative enough to have one.
|
|
298
366
|
const moved = !informativeHash(before) || !informativeHash(after) || before !== after;
|
|
299
|
-
|
|
367
|
+
const launched = ran(left);
|
|
368
|
+
if (!launched) skip('the screen actually changed before testing the stale ref',
|
|
300
369
|
'the second app never launched, so there was no screen change to test against');
|
|
301
370
|
else check(moved, 'the screen actually changed before testing the stale ref',
|
|
302
371
|
`${before.slice(0, 10)} -> ${after.slice(0, 10)}`
|
|
@@ -305,7 +374,16 @@ if (first) {
|
|
|
305
374
|
// reports "the stale-ref guard failed" for a device that never left the
|
|
306
375
|
// screen, which is a false accusation against the one layer this file exists
|
|
307
376
|
// to defend — and it is how this check has failed twice.
|
|
308
|
-
|
|
377
|
+
// A skip has to propagate. The precondition above reported NOT TESTED and this
|
|
378
|
+
// check ran anyway and failed — which is the harness doing to itself, one line
|
|
379
|
+
// later, exactly what `skip` was written to stop it doing. `moved` is true
|
|
380
|
+
// when a hash is too degenerate to have a say, and that is right for "did the
|
|
381
|
+
// screen change" and wrong as a licence to run a check whose setup is known
|
|
382
|
+
// not to have happened.
|
|
383
|
+
if (!launched) {
|
|
384
|
+
skip('a ref numbered on another screen refuses instead of tapping those coordinates',
|
|
385
|
+
'we never reached another screen, so there was nothing to refuse from');
|
|
386
|
+
} else if (moved) {
|
|
309
387
|
// This matched on prose twice and went red twice, both times for a refusal
|
|
310
388
|
// that was correct and better worded than the alternation knew — most
|
|
311
389
|
// recently `"Welcome to Reminders" is not on this screen`, which refuses
|