simframe 0.12.2 → 0.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -30,6 +30,41 @@ flows run in one call, screens the agent has seen before are answered from
30
30
  memory, and every answer is text with tap points in it. Nothing returns an
31
31
  image unless you ask for one.
32
32
 
33
+ ### It is also a study case, and the numbers are the point
34
+
35
+ Almost everything here was decided by a measurement rather than by an argument,
36
+ and several of those measurements **reversed a decision that had already
37
+ shipped**. They are written down in full, with what we expected beforehand, what
38
+ it cost to find out, and the mistakes made getting there — because a benchmark
39
+ that only records the winner teaches nothing.
40
+
41
+ | | |
42
+ | --- | --- |
43
+ | [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) | the record: N, median, p95, machine, every round |
44
+ | [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) | what we believed *before* measuring, and the fourteen times it was wrong |
45
+ | [`docs/DEFERRED.md`](docs/DEFERRED.md) | every known defect, open or closed, with the evidence |
46
+ | [`docs/ARTICLE.md`](docs/ARTICLE.md) | the argument the whole thing adds up to |
47
+
48
+ The headline results, all free to reuse:
49
+
50
+ - **A warm transition graph removes ~27% of an agent's model round trips** —
51
+ isolated with a three-pass field experiment, and corroborated to three points
52
+ by an independent estimate. Operator knowledge is worth twice that, and no
53
+ tool here carries it yet. ([§14](docs/EXPERIMENTS.md))
54
+ - **A bigger local judge is not a better one.** `qwen3:14b` scored *lower* than
55
+ `qwen3:8b` — 82% against 91% — while being 79% larger and 62% slower. A
56
+ one-line threshold on a number already computed beat all of them at **95% and
57
+ zero latency**.
58
+ - **Every interactive element came from the accessibility tree — 72 of 72 — and
59
+ 83% of them have no text at all.** Screen *recognition*, though, did not need
60
+ the tree once. The ladder splits there: semantics need a tree, identity needs
61
+ pixels.
62
+ - **Capture is ~31× faster than a screenshot** for the same work, and the
63
+ capture primitive itself about a thousand times faster.
64
+ - **Three independent reporters converged on one thing**: a refusal is cheap and
65
+ a confident wrong answer is expensive. Every serious bug they found was a
66
+ component reporting more certainty than it had.
67
+
33
68
  ## What changed, measured
34
69
 
35
70
  Same four-tab navigation flow, on a real production app:
@@ -177,13 +212,30 @@ was otherwise impossible.
177
212
 
178
213
  **What is measured and what is not.** The ranker: 5 of 6 top-1 on hand-written
179
214
  cases, median 564 ms warm, and on a real exploration it went to the right region
180
- in two steps where reading order wandered into version strings. The supervisor:
181
- correct on four real batch-killers once the plan briefed it, 689–751 ms warm —
182
- **on a bench, not in the field.** `ax-first` made no measurable difference to how
183
- an agent drove a real app, with one small regression and one small win. Numbers
184
- and conditions are in [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md); the judgements,
185
- including a phase cancelled by its own measurement, are in
186
- [`docs/DECISIONS.md`](docs/DECISIONS.md).
215
+ in two steps where reading order wandered into version strings. `ax-first` made
216
+ no measurable difference to how an agent drove a real app, with one small
217
+ regression and one small win.
218
+
219
+ **The supervisor, on 22 labelled failures from a seeded React Native app** —
220
+ and the number to judge it by is round trips, not accuracy:
221
+
222
+ | | |
223
+ |---|---|
224
+ | handled locally, no round trip | **9 of 22** |
225
+ | escalated to Claude | 13 of 22 |
226
+ | a model round trip, measured in the field | 10–16 s |
227
+
228
+ So roughly 90–145 s saved on that population. Accuracy was 77 / 82 / 86% across
229
+ three runs of the *same* questions — it is **not deterministic**, and every
230
+ earlier single-run figure in this project carried that spread without reporting
231
+ it. All of its errors were `wait` where `stop` was right, which is the cheap
232
+ direction: a wrong `wait` costs a settle and a re-run, and `stop` already means
233
+ "hand back to Claude with the unattempted steps" rather than "give up".
234
+
235
+ **Still a bench, not the field.** These are fixtures we designed, on a testbed we
236
+ built, labelled by the person who then scored them. Numbers and conditions are in
237
+ [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md); the judgements, including a phase
238
+ cancelled by its own measurement, are in [`docs/DECISIONS.md`](docs/DECISIONS.md).
187
239
 
188
240
  **And the measurements that changed our minds** are in
189
241
  [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md), with what we expected beforehand
@@ -212,6 +264,15 @@ The Ollama arm is an **experiment, not a recommendation**: off unless named,
212
264
  no weights shipped, no dependency added, and every arm reads the same briefing
213
265
  out of `native/supervise.swift` so no arm is answering a different question.
214
266
 
267
+ **Two further results, both negative, both worth more than the table.** A
268
+ *cascade* — free rule first, model where it is unsure, Claude after that — was
269
+ worse at every abstention band we tried (95% for the rule alone; 91%, 86%, 82%
270
+ as more was handed to the model). And giving the model a fourth word, `abstain`,
271
+ cost the Apple arm about a third of its accuracy (77/82/86% → 45/50/55%) **while
272
+ it never used the word once**. One added paragraph, nothing else changed. So the
273
+ fourth word ships off, and "add an abstain token" became "find a judge that will
274
+ use one".
275
+
215
276
  The supervisor's whole vocabulary is three words on purpose. It cannot invent a
216
277
  step, skip one, substitute a target or continue past an unexpected screen — not
217
278
  because a threshold forbids it but because those are not answers it can give.
@@ -345,6 +406,7 @@ steer the model is a tool surface the model uses wrong.
345
406
  | `sim_wait` | Waits for the screen to change *and then* settle. |
346
407
  | `sim_look` | **The only tool that returns an image**, capped at 1024 px. For layout, colour, spacing — questions text cannot answer. |
347
408
  | `sim_recall` · `sim_strip` | Look backwards: a text timeline of what happened, or recent frames tiled into one image. |
409
+ | `sim_storage` | **What the app believes**, as opposed to what it drew: its `UserDefaults` and, for React Native, its `AsyncStorage`. Reads the data container off disk, so it answers on a device that is **not running**. |
348
410
  | `sim_capture` · `sim_devices` | Manage capture loops; list simulators. |
349
411
 
350
412
  ### What the screen looks like as text
@@ -524,6 +586,20 @@ simframe screens # what this device has learned
524
586
  simframe goto invoices # walk there, verifying every step
525
587
  ```
526
588
 
589
+ And when the screen and the behaviour disagree, the question is usually not
590
+ about the screen at all:
591
+
592
+ ```bash
593
+ simframe storage # apps with a data container
594
+ simframe storage com.example.myapp # what that app saved
595
+ ```
596
+
597
+ `sim_ui` says what is drawn; `sim_storage` says what the app believes. It reads
598
+ the data container straight off the host filesystem, which means it works on a
599
+ device that is **shut down** — `simctl` cannot do this at all, on any of its own
600
+ paths, once a device stops running.
601
+
602
+
527
603
  Measured on a four-tab tour, `goto` plans and walks three-step routes with every
528
604
  step verified and no model call. It fails rather than guesses: an unknown
529
605
  destination, a query matching two screens equally, or no path of known edges all
@@ -538,6 +614,29 @@ simframe flow save checkout ./checkout.json
538
614
  simframe flow run checkout
539
615
  ```
540
616
 
617
+ ### What the memory is worth, isolated
618
+
619
+ Three agent sessions drove the same task family on the same production app. The
620
+ third existed only to hold a variable still: pass 2 beat pass 1 by twenty-one
621
+ model round trips, but it had a warm graph *and* an operator who had already
622
+ driven the app once, so the gap could belong to either.
623
+
624
+ | pass | operator | graph | round trips |
625
+ | --- | --- | --- | --- |
626
+ | 1 | fresh | cold | **33** |
627
+ | 3 | fresh | **warm** | **24** |
628
+ | 2 | experienced | warm | **12** |
629
+
630
+ **The graph is worth ~27%** of the round trips. Operator knowledge is worth the
631
+ rest — half of everything left after the graph had taken its share. A second
632
+ reporter, unable to isolate the graph, derived **~23%** from escalation rate per
633
+ step without knowing that number.
634
+
635
+ One app, one operator per pass, and round trips are not seconds. What it settles
636
+ is the direction: the memory is real, it is the smaller half, and the larger half
637
+ is what an operator learns and no tool here carries between sessions yet. Full
638
+ working in [`docs/EXPERIMENTS.md`](docs/EXPERIMENTS.md) §14.
639
+
541
640
  ## Does this work on *your* app?
542
641
 
543
642
  Nothing in simframe is written for a particular app. What varies between apps is
@@ -553,9 +652,27 @@ Run `simframe ui` on any screen to see exactly what simframe can see, with each
553
652
  target marked `ax` or `ocr`. If something you can read is not listed, that is a
554
653
  bug worth reporting.
555
654
 
655
+ A control the app **declares but never names** — an icon-only overflow menu, a
656
+ back chevron — is listed with its coordinates and tappable by `#ref` rather than
657
+ by name, and the map counts how many such controls are on the screen. A view the
658
+ app never declared accessible at all is invisible to any accessibility tree, ours
659
+ included, and the same line says so: a count of zero on a screen that has one
660
+ would be the more expensive answer.
661
+
662
+ **What a settle can and cannot see.** Stillness is decided from a mean over a
663
+ grid of the frame, and a small animation does not move a mean. Measured on a
664
+ screen built to never settle — a spinner, frames every 77–95 ms — the mean
665
+ difference was **0.00196** against a 0.004 threshold while one cell moved by
666
+ **0.0275**, and `stableForMs` reported **79 seconds** of stillness. So a settle
667
+ can return satisfied while part of the screen is still moving, and it now says
668
+ so: `settled after 63ms (a 13x13 region is still animating)`. A settle that
669
+ times out names where the movement is instead. Refusing to settle on any
670
+ animation would be right for a spinner and wrong for a blinking cursor, so the
671
+ disagreement is reported rather than resolved by a guessed threshold.
672
+
556
673
  Two honest caveats. OCR reads **text**, so a purely graphical icon with no label
557
- is invisible to both paths — use `sim_ui` to get its coordinates from the tree,
558
- or tap by position. And the confirm-button vocabulary (`APPLY`, `OK`, `SAVE`,
674
+ is invisible to both paths — the tree still gives you its coordinates, and
675
+ `#ref` still taps it. And the confirm-button vocabulary (`APPLY`, `OK`, `SAVE`,
559
676
  `DONE`…) is English; a localised UI needs those words extended.
560
677
 
561
678
  ## Measured
@@ -650,6 +767,18 @@ accepted as the name that loop used to have.
650
767
  a support banner was swallowing — a banner that was *in the element list the
651
768
  same call printed*. Nothing said "there is something at y≈753 and you started
652
769
  at y=750".
770
+ - **A control half off the edge is offered, clamped, and says so.** A filter
771
+ chip with 29 pt of itself visible used to be dropped — the off-screen test
772
+ reads the *centre* — and OCR's reading of that same sliver, a box labelled
773
+ `Flc`, was printed in its place. So the map did not merely omit a control; it
774
+ offered a meaningless name for it at an ordinary-looking coordinate. Now:
775
+
776
+ ```
777
+ #6 button 387,191 Flowering (partly off-screen — the coordinate is the middle of the visible part)
778
+ ```
779
+
780
+ `scrollTo` is deliberately unchanged: a clipped element is still "not in
781
+ view", so it still scrolls rather than calling a sliver good enough.
653
782
 
654
783
  ## CLI
655
784
 
@@ -13,6 +13,10 @@ public enum Motion {
13
13
 
14
14
  /// Below this mean absolute difference, two frames are the same picture.
15
15
  public static let stillThreshold = 0.004
16
+
17
+ /// How many moving cells make an animation rather than sensor noise.
18
+ /// Set from the measurements in `state(history:now:)` below.
19
+ public static let minAnimatingCells = 12
16
20
  /// Frames that must agree before the screen counts as settled.
17
21
  ///
18
22
  /// Paired with a duration, because frame count alone is not a measure of
@@ -170,8 +174,35 @@ public enum Motion {
170
174
  let cellThreshold = 0.06
171
175
  let moving = union.filter { $0 > cellThreshold }.count
172
176
  let fraction = Double(moving) / Double(union.count)
173
- // Small and persistent, rather than a screen changing.
174
- if fraction > 0 && fraction < 0.06 {
177
+ // Small and persistent, rather than a screen changing — and big
178
+ // enough to be something.
179
+ //
180
+ // `fraction > 0` meant one cell of 4,608 counted as an animation,
181
+ // and measured on a device that fires constantly on screens where
182
+ // nothing is happening. A field report put it exactly right: a
183
+ // warning that is usually wrong trains the reader to ignore the one
184
+ // that matters.
185
+ //
186
+ // Measured on this device, 2026-09-14, which is what
187
+ // `minAnimatingCells` is set from:
188
+ //
189
+ // static home screen, nothing moving 1-4 cells, on 54% of frames
190
+ // a blinking text caret **2 cells** (1x2), on 100%
191
+ // the testbed's spinner 169 cells (13x13)
192
+ // Maps launching, median 656 cells (41x16)
193
+ // a real spinner, from the field 1,364 cells (44x31)
194
+ //
195
+ // Caret and noise sit at 1-4; the smallest real animation seen is
196
+ // 169. A 42x gap with nothing in it, so the threshold is not
197
+ // delicate — 12 is three times the worst noise and an order of
198
+ // magnitude below the weakest signal.
199
+ //
200
+ // The caret is the case this is really for, and it is why the
201
+ // question could not be settled by reasoning: a caret must never
202
+ // stop a screen from settling, and before this it did — 72 frames
203
+ // out of 72 with `settled: false` on a screen holding nothing but a
204
+ // text cursor.
205
+ if moving >= Self.minAnimatingCells && fraction < 0.06 {
175
206
  animating = boundingBox(of: union, threshold: cellThreshold)
176
207
  }
177
208
  }
@@ -40,7 +40,40 @@ struct Judgement {
40
40
  var decision: Decision
41
41
  }
42
42
 
43
+ /**
44
+ * The same question with a fourth answer: "I cannot tell".
45
+ *
46
+ * A second type rather than a fourth case on `Decision`, because the two have
47
+ * to be *comparable*. `@Generable` fixes the answer space at compile time, so a
48
+ * binary that can only be asked one way cannot measure whether the token helps
49
+ * or merely moves the errors around — and "does the model abstain on the ones
50
+ * it gets wrong, or at random" is the whole question. One binary, both
51
+ * vocabularies, chosen per request.
52
+ *
53
+ * `abstain` is not a new capability and cannot become one. It means "behave as
54
+ * if there is no supervisor", which is exactly the `null` the caller already
55
+ * handles on every failure path. The answer space stays the safety property:
56
+ * this strictly shrinks what the model can cause to happen.
57
+ */
58
+ @available(macOS 26.0, *)
59
+ @Generable
60
+ enum CautiousDecision: String {
61
+ case wait
62
+ case retry
63
+ case stop
64
+ case abstain
65
+ }
66
+
67
+ @available(macOS 26.0, *)
68
+ @Generable
69
+ struct CautiousJudgement {
70
+ @Guide(description: "wait if the screen is still arriving, retry if the same step should be attempted again, stop if nothing further can work, abstain if you cannot tell from what you were given")
71
+ var decision: CautiousDecision
72
+ }
73
+
43
74
  struct Situation: Decodable {
75
+ /// Ask with the fourth word available. Absent means the shipped three.
76
+ let mayAbstain: Bool?
44
77
  let goal: String?
45
78
  let step: String
46
79
  let expected: String?
@@ -121,6 +154,19 @@ func serve() async {
121
154
  are not being asked what to do, only whether this can proceed. Answer with \
122
155
  the decision alone.
123
156
  """
157
+
158
+ // The same brief with the fourth word, built from the first so the two
159
+ // cannot drift. Everything above applies; this adds one rule and changes
160
+ // nothing else, which is what makes the comparison mean something.
161
+ let abstainInstructions = instructions + "\n\n" + """
162
+ There is a fourth answer: abstain. Use it when the evidence you were \
163
+ given does not distinguish these cases — when you would be guessing. \
164
+ Abstaining is not a failure and it is not penalised: it hands the decision \
165
+ to someone with more context, which costs one step. A wrong stop abandons a \
166
+ plan that would have worked, and a wrong wait spends a timeout on a screen \
167
+ that will never change. Prefer abstain to either. Do not abstain merely \
168
+ because the situation is unusual; abstain because the evidence is absent.
169
+ """
124
170
  // Warm up front so the first real judgement does not pay model load —
125
171
  // measured at ~880ms against ~600ms warm.
126
172
  //
@@ -133,7 +179,17 @@ func serve() async {
133
179
  // The window, reported rather than assumed. 4,096 has been a documented
134
180
  // constant we repeated; since 26.4 it is queryable, so it is now read from
135
181
  // the model and handed to the caller, who prints it in `doctor`.
136
- emit(["ready": true, "contextSize": SystemLanguageModel.default.contextSize])
182
+ // Guarded at 26.4, which is where `contextSize` became queryable — the
183
+ // enclosing block only promises 26.0, so an SDK between the two would fail
184
+ // to *compile* and the supervisor would simply not exist on that machine,
185
+ // reported as "could not build the local supervisor". A peer round is
186
+ // exactly where that costs somebody an afternoon, and the window is a
187
+ // nicety: it is printed in `doctor` and nothing depends on it.
188
+ var hello: [String: Any] = ["ready": true]
189
+ if #available(macOS 26.4, *) {
190
+ hello["contextSize"] = SystemLanguageModel.default.contextSize
191
+ }
192
+ emit(hello)
137
193
  while let line = readLine(strippingNewline: true) {
138
194
  if line.isEmpty { continue }
139
195
  guard let data = line.data(using: .utf8),
@@ -157,8 +213,11 @@ func serve() async {
157
213
  }
158
214
  let started = Date()
159
215
  do {
160
- let session = LanguageModelSession(instructions: instructions)
161
- let out = try await session.respond(to: prompt, generating: Judgement.self)
216
+ let mayAbstain = s.mayAbstain ?? false
217
+ let session = LanguageModelSession(instructions: mayAbstain ? abstainInstructions : instructions)
218
+ let decision: String = mayAbstain
219
+ ? try await session.respond(to: prompt, generating: CautiousJudgement.self).content.decision.rawValue
220
+ : try await session.respond(to: prompt, generating: Judgement.self).content.decision.rawValue
162
221
  // No reason field, deliberately. Asked for one it confabulated in
163
222
  // every observed run: a correct `stop` justified as "screen is
164
223
  // elsewhere" when the screen was exactly where the plan expected,
@@ -170,7 +229,7 @@ func serve() async {
170
229
  // one."* What the caller gets instead is which rule or which model
171
230
  // answered, which is true by construction.
172
231
  emit([
173
- "decision": out.content.decision.rawValue,
232
+ "decision": decision,
174
233
  "ms": Int(Date().timeIntervalSince(started) * 1000),
175
234
  ])
176
235
  } catch let err as LanguageModelSession.GenerationError {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "simframe",
3
- "version": "0.12.2",
3
+ "version": "0.14.0",
4
4
  "mcpName": "io.github.lvlrSajjad/simframe",
5
5
  "description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
6
6
  "keywords": [
@@ -0,0 +1,185 @@
1
+ #!/usr/bin/env node
2
+ // Generate docs/ARTICLE.md from docs/agents-shouldnt-blink.html.
3
+ //
4
+ // node scripts/article-md.mjs # write it
5
+ // node scripts/article-md.mjs --check # fail if it is out of date
6
+ //
7
+ // **Why generated rather than kept in step by hand.** The two were written as a
8
+ // pair, with a note at the top of the Markdown saying it was the source of
9
+ // record and to edit both together. Then the page grew several sections the
10
+ // Markdown never got — the local supervisor, the capacity comparison — and the
11
+ // note stayed there being wrong, which is worse than no note: the next person
12
+ // follows it. Replacing a stale instruction with a stale *warning* would have
13
+ // been the same mistake one level up.
14
+ //
15
+ // So the page is the source and this is the projection. Two copies of a
16
+ // document cannot drift when one of them is derived, and `--check` in CI means
17
+ // they cannot drift silently even for a commit.
18
+ //
19
+ // This is not a general HTML-to-Markdown converter and should not become one.
20
+ // It understands exactly the constructs that article uses, and it *throws* on
21
+ // anything it does not recognise rather than dropping it — a converter that
22
+ // silently skips an unfamiliar tag produces a document that looks complete and
23
+ // is not, which is the failure mode this whole file exists to prevent.
24
+ import fs from 'node:fs';
25
+ import path from 'node:path';
26
+ import { fileURLToPath } from 'node:url';
27
+
28
+ const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
29
+ const SOURCE = path.join(ROOT, 'docs', 'agents-shouldnt-blink.html');
30
+ const TARGET = path.join(ROOT, 'docs', 'ARTICLE.md');
31
+
32
+ const ENTITIES = {
33
+ mdash: '—', ndash: '–', nbsp: ' ', amp: '&', lt: '<', gt: '>', quot: '"',
34
+ rsquo: '’', lsquo: '‘', ldquo: '“', rdquo: '”',
35
+ times: '×', rarr: '→', larr: '←', hellip: '…', plusmn: '±', deg: '°',
36
+ middot: '·', bull: '•', frac12: '½', le: '≤', ge: '≥', minus: '−', apos: "'",
37
+ };
38
+
39
+ /** Entities, then inline tags, then whitespace. Order matters. */
40
+ function inline(html) {
41
+ let out = html
42
+ // The faculty spans open a paragraph and name the thing it defines —
43
+ // "Eyes", "Hand", "Memory" — with no separator in the HTML because the
44
+ // stylesheet supplies one.
45
+ .replace(/<span class="faculty">([\s\S]*?)<\/span>/g, (_, t) => `<strong>${t}</strong> — `)
46
+ // An abbreviation carries its expansion in the title attribute, which a
47
+ // reader of the Markdown has no way to hover over. Keep both.
48
+ .replace(/<abbr title="([^"]*)"[^>]*>([\s\S]*?)<\/abbr>/g, (_, full, t) => `${stripTags(t)} (${full})`)
49
+ .replace(/<a\b[^>]*href="([^"]*)"[^>]*>([\s\S]*?)<\/a>/g, (_, href, text) => `[${inline(text)}](${href})`)
50
+ .replace(/<code>([\s\S]*?)<\/code>/g, (_, t) => `\`${stripTags(t)}\``)
51
+ .replace(/<(?:b|strong)>([\s\S]*?)<\/(?:b|strong)>/g, (_, t) => `**${inline(t)}**`)
52
+ .replace(/<(?:i|em)>([\s\S]*?)<\/(?:i|em)>/g, (_, t) => `*${inline(t)}*`)
53
+ .replace(/<br\s*\/?>/g, ' \n');
54
+ const left = out.match(/<(?!\/)[a-z]/i);
55
+ if (left) throw new Error(`unhandled inline tag near: ${out.slice(Math.max(0, left.index - 40), left.index + 60)}`);
56
+ return decode(out).replace(/\s+/g, ' ').trim();
57
+ }
58
+
59
+ const decode = (s) => s
60
+ .replace(/&#(\d+);/g, (_, n) => String.fromCodePoint(Number(n)))
61
+ .replace(/&([a-zA-Z0-9]+);/g, (m, name) => {
62
+ if (!(name in ENTITIES)) throw new Error(`unknown entity &${name};`);
63
+ return ENTITIES[name];
64
+ });
65
+
66
+ const stripTags = (s) => decode(s.replace(/<[^>]+>/g, '')).replace(/\s+/g, ' ').trim();
67
+
68
+ /** A table, as GitHub-flavoured Markdown. */
69
+ function table(html) {
70
+ const rows = [...html.matchAll(/<tr>([\s\S]*?)<\/tr>/g)].map((m) =>
71
+ [...m[1].matchAll(/<(th|td)\b[^>]*>([\s\S]*?)<\/\1>/g)].map((c) => inline(c[2])));
72
+ if (!rows.length) return [];
73
+ const head = /<thead>/.test(html) ? rows[0] : rows[0].map(() => ' ');
74
+ const body = /<thead>/.test(html) ? rows.slice(1) : rows;
75
+ const width = Math.max(...rows.map((r) => r.length));
76
+ const pad = (r) => [...r, ...Array(width - r.length).fill('')];
77
+ return [
78
+ `| ${pad(head).join(' | ')} |`,
79
+ `| ${Array(width).fill('---').join(' | ')} |`,
80
+ ...body.map((r) => `| ${pad(r).join(' | ')} |`),
81
+ ];
82
+ }
83
+
84
+ /** Everything inside one `<section>`, in order. */
85
+ function blocks(html) {
86
+ const out = [];
87
+ // One pass over the constructs the article actually uses. `lastIndex`
88
+ // walking rather than nested parsing, because the document is flat inside a
89
+ // section and a real parser here would be a dependency.
90
+ const re = new RegExp([
91
+ '<h([123])>([\\s\\S]*?)<\\/h\\1>',
92
+ '<p class="sec-label">([\\s\\S]*?)<\\/p>',
93
+ '<p class="panel-label">([\\s\\S]*?)<\\/p>',
94
+ '<span class="amend-label">([\\s\\S]*?)<\\/span>',
95
+ '<p\\b[^>]*>([\\s\\S]*?)<\\/p>',
96
+ '<table>([\\s\\S]*?)<\\/table>',
97
+ '<pre>([\\s\\S]*?)<\\/pre>',
98
+ '<li>([\\s\\S]*?)<\\/li>',
99
+ ].join('|'), 'g');
100
+ let m;
101
+ while ((m = re.exec(html)) !== null) {
102
+ const [, level, heading, sec, panel, amend, para, tbl, pre, li] = m;
103
+ if (heading != null) out.push({ kind: 'h', level: Number(level), text: inline(heading) });
104
+ else if (sec != null) out.push({ kind: 'kicker', text: inline(sec) });
105
+ else if (panel != null) out.push({ kind: 'label', text: inline(panel) });
106
+ else if (amend != null) out.push({ kind: 'label', text: inline(amend) });
107
+ else if (para != null) out.push({ kind: 'p', text: inline(para) });
108
+ else if (tbl != null) out.push({ kind: 'table', lines: table(tbl) });
109
+ else if (pre != null) out.push({ kind: 'pre', text: decode(stripTagsPreserving(pre)) });
110
+ else if (li != null) out.push({ kind: 'li', text: inline(li) });
111
+ }
112
+ return out;
113
+ }
114
+
115
+ /** `<pre>` keeps its newlines; only tags come out. */
116
+ const stripTagsPreserving = (s) => s.replace(/<[^>]+>/g, '').replace(/^\n/, '').replace(/\s+$/, '');
117
+
118
+ export function render(html) {
119
+ const title = stripTags(html.match(/<h1>([\s\S]*?)<\/h1>/)[1]);
120
+ const kicker = stripTags(html.match(/<p class="kicker">([\s\S]*?)<\/p>/)[1]);
121
+ const standfirst = html.match(/<p class="standfirst">([\s\S]*?)<\/p>/);
122
+
123
+ const lines = [
124
+ `# ${title}`,
125
+ '',
126
+ `*${kicker}*`,
127
+ '',
128
+ '> **This file is generated** from [`agents-shouldnt-blink.html`](agents-shouldnt-blink.html)',
129
+ '> by `scripts/article-md.mjs`. Edit the page, not this — a hand-kept copy of a',
130
+ '> document is a copy that drifts, and this one did.',
131
+ '>',
132
+ '> The measurements, with N, median, p95 and the mistakes made getting to each,',
133
+ '> are in [`BENCHMARKS.md`](BENCHMARKS.md); what we expected before measuring is',
134
+ '> in [`EXPERIMENTS.md`](EXPERIMENTS.md).',
135
+ '',
136
+ ];
137
+ if (standfirst) lines.push(inline(standfirst[1]), '');
138
+ lines.push('---', '');
139
+
140
+ for (const section of html.match(/<section>[\s\S]*?<\/section>/g) ?? []) {
141
+ let pendingKicker = null;
142
+ let inList = false;
143
+ for (const b of blocks(section)) {
144
+ if (b.kind !== 'li' && inList) { lines.push(''); inList = false; }
145
+ switch (b.kind) {
146
+ // The kicker precedes its heading on the page and reads as a label for
147
+ // it, so it is held until the heading arrives rather than emitted where
148
+ // it was found.
149
+ case 'kicker': pendingKicker = b.text; break;
150
+ case 'h':
151
+ if (pendingKicker) { lines.push(`*${pendingKicker}*`, ''); pendingKicker = null; }
152
+ // `<h1>` is the title, already emitted above, so the page's `<h2>`
153
+ // sections are the document's second level and not its third.
154
+ lines.push(`${'#'.repeat(b.level)} ${b.text}`, '');
155
+ break;
156
+ case 'label': lines.push(`**${b.text}**`, ''); break;
157
+ case 'p': lines.push(b.text, ''); break;
158
+ case 'li': lines.push(`- ${b.text}`); inList = true; break;
159
+ case 'table': lines.push(...b.lines, ''); break;
160
+ case 'pre': lines.push('```', b.text, '```', ''); break;
161
+ default: throw new Error(`unhandled block ${b.kind}`);
162
+ }
163
+ }
164
+ if (inList) lines.push('');
165
+ }
166
+
167
+ return `${lines.join('\n').replace(/\n{3,}/g, '\n\n').trim()}\n`;
168
+ }
169
+
170
+ const html = fs.readFileSync(SOURCE, 'utf8');
171
+ const md = render(html);
172
+
173
+ if (process.argv.includes('--check')) {
174
+ const current = fs.existsSync(TARGET) ? fs.readFileSync(TARGET, 'utf8') : '';
175
+ if (current === md) {
176
+ console.log(`docs/ARTICLE.md is in step with the page (${md.split('\n').length} lines)`);
177
+ process.exit(0);
178
+ }
179
+ console.error('docs/ARTICLE.md is out of date with docs/agents-shouldnt-blink.html.');
180
+ console.error('Run: node scripts/article-md.mjs');
181
+ process.exit(1);
182
+ }
183
+
184
+ fs.writeFileSync(TARGET, md);
185
+ console.log(`wrote docs/ARTICLE.md — ${md.split('\n').length} lines from ${SOURCE.split('/').pop()}`);
@@ -0,0 +1,82 @@
1
+ #!/usr/bin/env node
2
+ // Run a CI step, and tell a sick simulator apart from a failing check.
3
+ //
4
+ // node scripts/ci-device-guard.mjs <udid> -- <command> [args...]
5
+ //
6
+ // **Why this exists, with the number that justifies it.** Over the last 25 CI
7
+ // runs the `integration` job failed 17 times and passed 5. Eight of the nine
8
+ // most recent failures name a device-state condition in their own output —
9
+ // `NSPOSIXErrorDomain code=60`, `the display produced no frame in 60s`, `the
10
+ // second app never launched` — and one of those runs was a **docs-only commit**
11
+ // that changed a single markdown file. Two runs of byte-identical code failed at
12
+ // two different steps.
13
+ //
14
+ // So the job has been answering two questions at once — *does simframe work* and
15
+ // *did this hosted simulator survive twenty minutes* — and the second dominates.
16
+ // A red build that is usually the device is the cry-wolf failure this repo keeps
17
+ // writing items about, and it cost two days of re-reading logs to learn nothing.
18
+ //
19
+ // This does not paper over failures. It distinguishes them: a named device
20
+ // condition gets the cure this project already ships (`simframe revive`) and one
21
+ // retry, exactly as DEFERRED 126 says to; anything else fails on the spot,
22
+ // untouched. The classification is written to the job summary so the *rate*
23
+ // becomes visible instead of arguable.
24
+ import { spawn } from 'node:child_process';
25
+ import fs from 'node:fs';
26
+
27
+ /** Conditions that are the simulator, not the code. Each seen in a real run. */
28
+ const DEVICE_STATE = [
29
+ [/NSPOSIXErrorDomain.*code=?\s*60|Operation timed out/i, 'simctl stopped answering (NSPOSIXErrorDomain 60)'],
30
+ [/did not produce a frame|produced no frame in \d+s/i, 'the daemon is up and the display renders nothing'],
31
+ [/Timeout waiting for screen surfaces|display surface is not answering|display surface could not be read/i, 'the display surface is wedged'],
32
+ [/no frames buffered|capture is wedged/i, 'capture stopped'],
33
+ [/the second app never launched|could not be dispatched/i, 'an app would not launch'],
34
+ ];
35
+
36
+ const udid = process.argv[2];
37
+ const sep = process.argv.indexOf('--');
38
+ if (!udid || sep < 0) {
39
+ console.error('usage: ci-device-guard.mjs <udid> -- <command> [args...]');
40
+ process.exit(2);
41
+ }
42
+ const cmd = process.argv.slice(sep + 1);
43
+
44
+ function run(argv, { capture = true } = {}) {
45
+ return new Promise((resolve) => {
46
+ const p = spawn(argv[0], argv.slice(1), { stdio: capture ? ['inherit', 'pipe', 'pipe'] : 'inherit' });
47
+ let out = '';
48
+ p.stdout?.on('data', (d) => { out += d; process.stdout.write(d); });
49
+ p.stderr?.on('data', (d) => { out += d; process.stderr.write(d); });
50
+ p.on('close', (code) => resolve({ code, out }));
51
+ });
52
+ }
53
+
54
+ const summary = (line) => {
55
+ const f = process.env.GITHUB_STEP_SUMMARY;
56
+ if (f) { try { fs.appendFileSync(f, `${line}\n`); } catch { /* summaries are a nicety */ } }
57
+ };
58
+
59
+ const deviceCause = (text) => DEVICE_STATE.find(([re]) => re.test(text))?.[1] ?? null;
60
+
61
+ const first = await run(cmd);
62
+ if (first.code === 0) process.exit(0);
63
+
64
+ const cause = deviceCause(first.out);
65
+ if (!cause) {
66
+ console.error(`\n (this step failed on its merits, not on the device — not retrying)`);
67
+ summary(`- \`${cmd.join(' ')}\` — **check failed** (exit ${first.code})`);
68
+ process.exit(first.code ?? 1);
69
+ }
70
+
71
+ console.error(`\n (${cause} — DEFERRED 126. Reviving once and running again.)`);
72
+ await run(['node', 'src/cli.js', 'revive', `--device=${udid}`], { capture: false });
73
+ const second = await run(cmd);
74
+ if (second.code === 0) {
75
+ summary(`- \`${cmd.join(' ')}\` — passed after one revive (${cause})`);
76
+ process.exit(0);
77
+ }
78
+ // Twice in a row, on a condition we know the cure for. Reported as what it is.
79
+ const again = deviceCause(second.out);
80
+ summary(`- \`${cmd.join(' ')}\` — **${again ? 'device unavailable' : 'check failed'}** after a revive${again ? ` (${again})` : ''}`);
81
+ if (again) console.error(`\nFAIL the simulator is still in a bad state after a revive: ${again}`);
82
+ process.exit(second.code ?? 1);