simframe 0.16.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -172,6 +172,18 @@ selector that did not *resolve* and nothing else, because retrying from a screen
172
172
  you did not expect to be on is a second guess. A destructive-looking label is
173
173
  refused as a substitute even if you list it.
174
174
 
175
+ **Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
176
+ "optional": true}` is skipped when nothing matches and runs normally when
177
+ something does. It exists because the opposite cost real time: a batch that
178
+ included a dismissal for a first-launch sheet lost **six correct steps** on the
179
+ next run because the sheet *did not appear*. A nag screen, a permission prompt,
180
+ a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
181
+ is the expensive outcome — a call per step instead of a call per flow. Only a
182
+ selector that resolved to **nothing** is absorbed: a target that is on screen
183
+ twice is ambiguous, not absent, and must still verify. Skipped steps are
184
+ reported, because "the sheet was gone" and "the sheet was dismissed" are
185
+ different facts.
186
+
175
187
  **`{"seek": "change username", "budget": 6}`** opens containers, checks, and
176
188
  comes back, depth first, inside a hard budget. It **acts** — opening a door
177
189
  changes state — and it refuses to open anything that commits, abandons or
@@ -400,7 +412,7 @@ steer the model is a tool surface the model uses wrong.
400
412
  | `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
401
413
  | `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
402
414
  | `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
403
- | `sim_flow_run` | Replay a flow that verified end to end. |
415
+ | `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
404
416
  | `sim_find` | Resolve an intent to one control, without acting on it. |
405
417
  | `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
406
418
  | `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
@@ -710,16 +722,35 @@ feels, and it took two field reports to see it. Per-step wall clock is
710
722
  (model round trip + simframe work × n) / n for n steps in one call
711
723
  ```
712
724
 
713
- | | per step |
714
- | --- | --- |
715
- | a human tester, measured | **1.95 s** |
716
- | **a saved flow replayed — zero model calls** | **1.98 s** |
717
- | simframe's own work inside a batch | **~1.7 s** |
718
- | a batch of 4 steps, one model call | ~6.7 s |
719
- | a batch of 2 steps | ~11.7 s |
720
- | one model call per step | ~21.7 s |
721
-
722
- **A replayed flow runs at human speed**, and simframe's own work already does.
725
+ **Compare like with like, which this table used not to.** A human tester's
726
+ 1.95 s per step is the *whole* loop — look at the screen, decide what to do,
727
+ do it. Any row that does not include a decision is not comparable to it.
728
+
729
+ | | perceives | decides | acts | per step |
730
+ | --- | --- | --- | --- | --- |
731
+ | **a human tester, measured** | yes | yes | yes | **1.95 s** |
732
+ | **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
733
+ | simframe, batch of 4 | yes | yes | yes | ~6.7 s |
734
+ | simframe, one model call per step | yes | yes | yes | ~21.7 s |
735
+ | — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
736
+ | — a saved flow replayed | yes | **no** | yes | 1.98 s |
737
+
738
+ **The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
739
+ a human**, and about eleven times when a failure forces one model call per step.
740
+
741
+ The two indented rows are the ones this README used to lead with, and both are
742
+ category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
743
+ taken out — the thinking is the model round trip, which is most of the clock.
744
+ And a replay *decides nothing*: it is a recording being played back, so its fair
745
+ counterpart is a human repeating a flow they have memorised, who would be well
746
+ under 1.95 s. Replay against a human working something out for the first time is
747
+ a rehearsal measured against a first attempt.
748
+
749
+ What the two rows do establish, and it is the finding that reordered this
750
+ project: **there is nothing left to win inside the engine.** 1.7 s is small
751
+ beside a 20 s round trip, so making perception or input faster buys single-digit
752
+ percentages. The only variable that matters is `n` — how many steps one decision
753
+ covers. Every improvement here has come from raising it, not from faster code.
723
754
  A field report put the split at **34% simframe, 60% agent round trips** over
724
755
  462 s of wall clock — the tester's *"30+ seconds between each step"* was
725
756
  accurate and was not simframe. So there is nothing left to win inside the
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "simframe",
3
- "version": "0.16.0",
3
+ "version": "0.17.0",
4
4
  "mcpName": "io.github.lvlrSajjad/simframe",
5
5
  "description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
6
6
  "keywords": [
@@ -646,12 +646,25 @@ console.log('\n--- what may be saved, and what may not ---');
646
646
  // evidence for it. Checked against whatever the run actually produced rather
647
647
  // than assuming a shape, as before.
648
648
  const attempt = await jsonRetry(['do', LOOP, `--save=${FLOW_NAME}`], { allowFail: true });
649
+ //
650
+ // Three outcomes, not two. The missing one is a step that ran and *failed*: a
651
+ // failing step stops the batch, so a failure on the last step leaves
652
+ // `ranSteps === steps.length` and the run reads as complete. A peer watched
653
+ // `FLOW FAILED — 4 ok, 1 failed (of 5)` save itself. This branch is also the
654
+ // one that goes stale: yesterday the save contract was replaced, the unit test
655
+ // was updated, and this file was left asserting the old rule — so the three
656
+ // cases are enumerated here explicitly rather than left as an `else`.
649
657
  const contradicted = attempt.results.some((r) => /^unexpected/.test(r.verification?.verdict ?? ''));
658
+ const anyFailed = attempt.results.some((r) => !r.ok);
650
659
  const clean = attempt.results.every((r) => !r.verification || r.verification.verdict === 'ok');
651
660
  if (contradicted) {
652
661
  check(attempt.saved?.ok === false && attempt.saved?.reason === 'contradicted-steps',
653
662
  'a flow with a contradicted step is refused, not quietly saved',
654
663
  `${attempt.saved?.reason} (${(attempt.saved?.verdicts ?? []).join(', ')})`);
664
+ } else if (anyFailed) {
665
+ check(attempt.saved?.ok === false && attempt.saved?.reason === 'failed-steps',
666
+ 'a flow with a step that ran and failed is refused, not quietly saved',
667
+ `${attempt.saved?.reason} (failed step ${(attempt.saved?.failed ?? []).join(', ')})`);
655
668
  } else {
656
669
  check(attempt.saved?.ok === true, 'a first traversal is recordable',
657
670
  `${attempt.saved?.steps} steps`);
@@ -676,12 +689,21 @@ if (check(forced.saved?.ok === true, 'and --force saves it anyway', `${forced.sa
676
689
  // state nothing ever leaves: a clean replay is the confirmation a first
677
690
  // traversal could not give. Only asserted when the replay actually ran to
678
691
  // the end — a partial replay promotes nothing, deliberately.
692
+ const after = await jsonRetry(['flow', 'list']);
693
+ const entry = after.find((f) => f.name === FLOW_NAME);
679
694
  if (replayed.ok) {
680
- const after = await jsonRetry(['flow', 'list']);
681
- const entry = after.find((f) => f.name === FLOW_NAME);
682
695
  check(entry && !entry.provisional,
683
696
  'and a clean replay confirms a provisional flow',
684
697
  `provisional=${entry?.provisional ?? 'gone'}`);
698
+ } else {
699
+ // The other direction, and it is the one that was silently wrong: promotion
700
+ // keyed on `ranSteps === steps.length`, which a failure on the *last* step
701
+ // satisfies. Both of a peer's saved flows were marked confirmed by replays
702
+ // that failed. A confirmation that a failing replay can grant is not a
703
+ // confirmation, so the negative case has to be checked too.
704
+ check(entry?.provisional === true,
705
+ 'and a replay that failed does NOT confirm the flow it just disproved',
706
+ `ok=${replayed.ok}, provisional=${entry?.provisional}`);
685
707
  }
686
708
  const unknown = await cli(['flow', 'run', 'no-such-flow'], { expectFail: true });
687
709
  check(/no flow/i.test(unknown), 'an unknown flow name is refused with what is known');
package/src/actions.js CHANGED
@@ -346,6 +346,16 @@ export async function runScript(
346
346
  // the caller returns to Claude is rendered from this, so describing the end
347
347
  // state costs nothing beyond the verification pass the flow already ran.
348
348
  let endScreen = null;
349
+ // Where this run began, so a flow saved from it can say what it assumes.
350
+ //
351
+ // Null when the first step does not depend on where we were — see
352
+ // `resetsTheScreen`.
353
+ //
354
+ // It was recorded as `null` on every flow 0.16.0 saved, which a peer caught:
355
+ // "replay never checks it is starting where it was recorded". Reading it is
356
+ // free — the first step computes the identity anyway — and it is a fact the
357
+ // recording had and discarded.
358
+ let startScreen = null;
349
359
  // At most one recovery per run. Pressing home while already on the springboard
350
360
  // moves nothing and is not a failure, so an unbounded retry would rebuild the
351
361
  // session and press again on every such step for no reason.
@@ -393,6 +403,9 @@ export async function runScript(
393
403
  ? (carriedScreen ?? await api.screenIdentity(deviceQuery, { options, settleMs: stableMs, timeoutMs, confirmNovel }))
394
404
  : null;
395
405
  carriedScreen = null;
406
+ if (i === 0 && !resetsTheScreen(step)) {
407
+ startScreen = beforeScreen ? { hash: beforeScreen.hash ?? null, tokens: beforeScreen.tokens ?? null } : null;
408
+ }
396
409
  // What this action did last time it was taken here, if ever.
397
410
  const prediction = verify && beforeScreen?.hash ? graph.predict(udid, beforeScreen, step) : null;
398
411
  try {
@@ -945,6 +958,33 @@ export async function runScript(
945
958
  break;
946
959
  }
947
960
  } catch (err) {
961
+ // A step that was allowed not to be there, and is not there.
962
+ //
963
+ // `or:` supplies alternative *selectors*; nothing could express "skip
964
+ // this if absent". So a batch crossing a first-launch nag, a permission
965
+ // prompt or a "What's New" could not be batched at all — the interstitial
966
+ // appears on run one and not on run two, the step fails, and every
967
+ // remaining step is discarded. A peer lost six correct steps because a
968
+ // sheet *did not* appear, and on the framing that every hard-fail
969
+ // dropping a caller back to single-stepping is a latency bug, that is the
970
+ // most expensive shape there is: ~20 s per step, for the rest of the plan.
971
+ //
972
+ // Deliberately narrow. Only `unknown_screen` — the selector resolved to
973
+ // nothing — absorbs. `ambiguous_intent` means the target *is* there,
974
+ // twice, and a step that is present must still verify; so must one that
975
+ // resolved and then failed. "Skip if absent" must not become "tap
976
+ // whatever is there".
977
+ if (step.optional && didNotResolve(err)) {
978
+ results.push({
979
+ index: i,
980
+ action: step.action,
981
+ ok: true,
982
+ skipped: true,
983
+ ms: Date.now() - stepStart,
984
+ detail: `skipped — optional, and ${JSON.stringify(String(goalOf(step) ?? step.action))} is not on this screen`,
985
+ });
986
+ continue;
987
+ }
948
988
  results.push({ index: i, action: step.action, ok: false, ms: Date.now() - stepStart, error: err.message });
949
989
  const why = metrics.reasonForStepError(step, err);
950
990
  noteEscalation({
@@ -1020,6 +1060,7 @@ export async function runScript(
1020
1060
  // navigate.saveFlow without the caller reassembling what it just ran.
1021
1061
  steps,
1022
1062
  flowId,
1063
+ startScreen,
1023
1064
  endScreen,
1024
1065
  results,
1025
1066
  ok: !failed,
@@ -1462,6 +1503,53 @@ export function alternativesFor(step) {
1462
1503
  return list.map((v) => (typeof v === 'string' ? v : v?.value ?? v?.target ?? v?.label)).filter(Boolean);
1463
1504
  }
1464
1505
 
1506
+ /**
1507
+ * Does this step make the screen we were on irrelevant?
1508
+ *
1509
+ * A flow that opens with `launch`, a deep link, or the home button does not
1510
+ * assume anything about where it starts — it goes and puts the device
1511
+ * somewhere. Recording a start screen for such a flow records where the device
1512
+ * happened to be beforehand, which is noise, and then reports a mismatch on
1513
+ * every replay that began anywhere else.
1514
+ *
1515
+ * Measured, not reasoned: a flow opening with `launch --relaunch` replayed
1516
+ * perfectly and still printed "recorded starting on 8292b488, replayed from
1517
+ * 39351dab", because the recording had followed a run that ended on Settings
1518
+ * root and the replay followed one that ended on About. Most flows open with a
1519
+ * launch, so that note would have fired on most correct replays — which is
1520
+ * item 175's defect exactly, in a feature added the same day 175 was written up.
1521
+ */
1522
+ export function resetsTheScreen(step = {}) {
1523
+ if (step.action === 'launch' || step.action === 'openUrl') return true;
1524
+ return step.action === 'button' && /^home$/i.test(String(step.value ?? '').trim());
1525
+ }
1526
+
1527
+ /**
1528
+ * Did this failure mean "the thing is not here", as opposed to any other way a
1529
+ * step can fail? The one question `optional` is allowed to ask.
1530
+ *
1531
+ * **Not `reason === 'unknown_screen'`, which is what this said for one build and
1532
+ * what made `optional` a no-op on every screen it was built for.** The tag at
1533
+ * the throw site is chosen by whether the *screen* was recognised, not by
1534
+ * whether the *target* was found: on a screen recalled from memory an absent
1535
+ * target is tagged `ambiguous_intent` too. So a conditional interstitial on a
1536
+ * well-known screen — the whole use case — never absorbed. Caught by running it
1537
+ * against a device; the unit test exercised this predicate and passed, because
1538
+ * it tested the predicate rather than the path.
1539
+ *
1540
+ * The distinction that actually exists is `ambiguous`, set only where something
1541
+ * was found *several times over*, and the comment at that throw site had
1542
+ * already written down why it matters: "more time cannot make a thing unique,
1543
+ * and it can make an absent thing arrive." Absent is the case `optional` may
1544
+ * skip; present-twice must still verify.
1545
+ */
1546
+ const RESOLVE_REASONS = new Set(['unknown_screen', 'ambiguous_intent']);
1547
+
1548
+ export function didNotResolve(err) {
1549
+ const tagged = metrics.escalationOf(err);
1550
+ return Boolean(tagged && RESOLVE_REASONS.has(tagged.reason) && !tagged.ambiguous);
1551
+ }
1552
+
1465
1553
  export function mayRetryAfter(err) {
1466
1554
  const tagged = metrics.escalationOf(err);
1467
1555
  return Boolean(tagged && RESOLVE_FAILURES.has(tagged.reason));
@@ -3246,13 +3334,23 @@ export function settleEvidence(w) {
3246
3334
  */
3247
3335
  export function flowSummary(res, { withTime = true } = {}) {
3248
3336
  const time = withTime && Number.isFinite(res.totalMs) ? ` in ${res.totalMs}ms` : '';
3249
- if (res.ok) return `flow completed — ${res.ranSteps}/${res.totalSteps} steps${time}`;
3337
+ // A step that was allowed not to be there and was not there did not run, and
3338
+ // saying so is the difference between "the nag screen was gone" and "the nag
3339
+ // screen was dismissed". Same lesson as the denominator above: a count that
3340
+ // quietly means two things is the defect.
3341
+ const skipped = (res.results ?? []).filter((r) => r.skipped).length;
3342
+ // After the time, in parentheses. Written inline before it, this read
3343
+ // "skipped as absent in 3412ms" — which says the skipping took 3.4 seconds.
3344
+ const skips = skipped
3345
+ ? ` (${skipped} optional step${skipped === 1 ? '' : 's'} skipped as absent)`
3346
+ : '';
3347
+ if (res.ok) return `flow completed — ${res.ranSteps}/${res.totalSteps} steps${time}${skips}`;
3250
3348
  const failed = (res.results ?? []).filter((r) => r.ok === false).length || 1;
3251
3349
  const worked = Math.max(0, res.ranSteps - failed);
3252
3350
  const unattempted = Math.max(0, res.totalSteps - res.ranSteps);
3253
3351
  return `FLOW FAILED — ${worked} ok, ${failed} failed`
3254
3352
  + (unattempted ? `, ${unattempted} not attempted` : '')
3255
- + ` (of ${res.totalSteps})${time}`;
3353
+ + ` (of ${res.totalSteps})${time}${skips}`;
3256
3354
  }
3257
3355
 
3258
3356
  /**
package/src/cli.js CHANGED
@@ -285,6 +285,20 @@ async function main() {
285
285
  // is a different machine and should say so rather than be guessed at.
286
286
  if (flags.readyTimeoutMs) options.readyTimeoutMs = num(flags.readyTimeoutMs);
287
287
 
288
+ // `--help` on any command, not only as the command.
289
+ //
290
+ // `simframe hpi --help` did the two worst things in sequence: without
291
+ // `--device` it demanded a device *before* printing help, and with one it
292
+ // ignored `--help` and ran the report. A flag that is ignored is worse than
293
+ // one that is rejected — it silently does something other than what was
294
+ // asked. (The general case, unrecognised flags being dropped everywhere
295
+ // rather than here, is filed separately; this is the one that was measured
296
+ // costing a round trip.)
297
+ if (flags.help || flags.h) {
298
+ process.stdout.write(USAGE);
299
+ return;
300
+ }
301
+
288
302
  switch (command) {
289
303
  case undefined:
290
304
  case '-h':
@@ -977,6 +991,9 @@ async function main() {
977
991
  return;
978
992
  }
979
993
  emit(flags, res, [
994
+ ...(res.startedElsewhere
995
+ ? [`NOTE: recorded starting on ${res.startedElsewhere.recorded}, replayed from ${res.startedElsewhere.here}`]
996
+ : []),
980
997
  ...(res.results ?? []).map(stepLine),
981
998
  actions.flowSummary(res, { withTime: false }),
982
999
  ]);
package/src/graph.js CHANGED
@@ -869,7 +869,7 @@ export const VERDICTS = ['ok', 'no-visible-change', 'unexpected-screen', 'unveri
869
869
  * that already exists do its job. A string still works and still means "exact
870
870
  * match only", which is right for a stored prediction that has no tokens.
871
871
  */
872
- function sameScreen(udid, a, b) {
872
+ export function sameScreen(udid, a, b) {
873
873
  const hashOf = (v) => (typeof v === 'string' ? v : v?.hash);
874
874
  const ha = hashOf(a);
875
875
  const hb = hashOf(b);
package/src/matching.js CHANGED
@@ -224,8 +224,9 @@ export function rank(targets, intent, { screen } = {}) {
224
224
 
225
225
  // **A distinctive fragment of one long name, when nothing else came close.**
226
226
  //
227
- // Reported from the field: `waitFor "6322594"` gave up after 20 s on a screen
228
- // whose own "Visible:" list printed `Work Order #6322594`. The reporter's
227
+ // Reported from the field: a `waitFor` on a seven-digit record number gave up
228
+ // after 20 s on a screen whose own "Visible:" list printed that number inside
229
+ // a heading — `Record #<digits>`, seventeen characters. The reporter's
229
230
  // guess was that `#` was being treated as significant, or that the matcher
230
231
  // was anchored. It is neither — the number scores **0.287** against the 0.45
231
232
  // floor, because the substring branch in `nameScore` scales by how much of
package/src/mcp.js CHANGED
@@ -134,7 +134,7 @@ const TOOLS = [
134
134
  steps: {
135
135
  type: 'array',
136
136
  description:
137
- 'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
137
+ 'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and "optional":true to any step that may legitimately have nothing to act on, which is how you cross a first-launch nag, a permission sheet or a "What\'s New" in one batch — {"tap":"Not Now","optional":true} is skipped when nothing matches and runs normally when it does, so the rest of the plan survives an interstitial that did not appear — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
138
138
  items: { type: 'object' },
139
139
  },
140
140
  autoSettle: {
@@ -1077,11 +1077,25 @@ async function doScript(target, args, options) {
1077
1077
  + ' seen before to compare against. Replay it once with sim_flow_run and it'
1078
1078
  + ' is confirmed. A replay costs no model calls.'
1079
1079
  : ' — confirmed; replay with sim_flow_run for zero model calls';
1080
+ // One sentence per reason, because "incomplete-run" and "failed-steps" ask
1081
+ // for opposite things from the caller and the bare reason names neither.
1082
+ const why = {
1083
+ 'contradicted-steps': 'a step landed somewhere memory says it should not have,'
1084
+ + ' so this route is not safe to replay. Check the flagged step before re-running it.',
1085
+ 'failed-steps': 'a step ran and failed, so replaying this would fail the same way'
1086
+ + ' every time. Fix the step and re-run the batch, then save.',
1087
+ 'incomplete-run': 'the run stopped before its last step, so the rest of the route'
1088
+ + ' was never observed. Re-run it to the end, then save.',
1089
+ }[saved.reason];
1090
+ // `verdicts` has a hole wherever a step never produced one — a failed step
1091
+ // has no verification — and joining those printed `(unverified, unverified, )`.
1092
+ const seen = (saved.verdicts ?? []).filter(Boolean);
1080
1093
  lines.push(
1081
1094
  saved.ok
1082
1095
  ? `saved as flow "${saved.name}" (${saved.steps} steps)${how}`
1083
1096
  : `NOT saved as "${args.saveAs}": ${saved.reason}`
1084
- + `${saved.verdicts ? ` (${saved.verdicts.join(', ')})` : ''}`,
1097
+ + `${why ? ` — ${why}` : ''}`
1098
+ + `${seen.length ? ` (verdicts: ${seen.join(', ')})` : ''}`,
1085
1099
  );
1086
1100
  }
1087
1101
  const escalated = (res.results ?? []).some((r) => metrics.ESCALATING_VERDICTS.has(r.verification?.verdict));
@@ -1167,6 +1181,14 @@ async function flowRun(target, args, options) {
1167
1181
  };
1168
1182
  }
1169
1183
  const lines = [`flow "${args.name}"`, ...stepLines(res)];
1184
+ // Said out loud or not said at all. A flow now records the screen it was
1185
+ // recorded on, and a mismatch that only exists in the return value is a fact
1186
+ // nobody reads — which is the same shape as the check that could not fail.
1187
+ if (res.startedElsewhere) {
1188
+ lines.push(`NOTE: this flow was recorded starting on screen ${res.startedElsewhere.recorded},`
1189
+ + ` and this replay started on ${res.startedElsewhere.here}. Not refused — content-driven`
1190
+ + ' screens legitimately change identity — but if the steps below fail to resolve, this is why.');
1191
+ }
1170
1192
  lines.push('', await mapFrom(target, options, res.endScreen, { verdictLine: verdictLineFor(res.results) }));
1171
1193
  return { content: [text(lines.join('\n'))], isError: !res.ok };
1172
1194
  }
package/src/navigate.js CHANGED
@@ -170,14 +170,31 @@ const CONTRADICTED = /^unexpected/;
170
170
  * and a flow that did not reach its own last step.
171
171
  */
172
172
  export function saveFlow(udid, name, script, { force = false } = {}) {
173
- const verdicts = (script.results ?? []).map((r) => r.verification?.verdict);
173
+ const results = script.results ?? [];
174
+ const verdicts = results.map((r) => r.verification?.verdict);
174
175
  const contradicted = verdicts.filter((v) => v && CONTRADICTED.test(v));
175
176
  const ranAll = script.ranSteps == null
176
177
  || script.steps == null
177
178
  || script.ranSteps >= (script.steps?.length ?? 0);
179
+ // A step that ran and failed is not a step that ran.
180
+ //
181
+ // `ranSteps` counts steps *attempted*, and a failing step stops the batch —
182
+ // so a run whose failure is on the **last** step has `ranSteps === steps.length`
183
+ // and passed the gate above. A peer watched `FLOW FAILED — 4 ok, 1 failed (of
184
+ // 5)` save itself as a flow, which is a flow guaranteed to fail forever. The
185
+ // predicate wanted "every step succeeded" and was written as "every step was
186
+ // reached"; on every run except one they are the same sentence.
187
+ //
188
+ // `!r.ok` rather than `r.ok === false` on purpose: a result that does not say
189
+ // it succeeded has not said it succeeded, and a gate that only catches an
190
+ // explicit `false` is one absent field away from being unable to fail.
191
+ const failedSteps = results.filter((r) => !r.ok);
178
192
  if (!force && contradicted.length) {
179
193
  return { ok: false, reason: 'contradicted-steps', verdicts };
180
194
  }
195
+ if (!force && failedSteps.length) {
196
+ return { ok: false, reason: 'failed-steps', verdicts, failed: failedSteps.map((r) => r.index ?? null) };
197
+ }
181
198
  if (!force && !ranAll) {
182
199
  return { ok: false, reason: 'incomplete-run', verdicts };
183
200
  }
@@ -249,6 +266,28 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
249
266
  // a human doing the same thing. `minSteps` comes from the flow definition or
250
267
  // stays null — the step count of a recorded route is not a claim about the
251
268
  // shortest one.
269
+ // Is this the screen the flow was recorded on?
270
+ //
271
+ // Reported, not refused, and the distinction is deliberate. A replay from the
272
+ // wrong screen is mostly self-limiting — the first selector does not resolve
273
+ // and the run stops one step in, and the destructive vocabulary is barred by
274
+ // the verify barrier either way. Refusing on a mismatch would put the exact
275
+ // failure mode of item 174 — content-driven screens fragmenting into several
276
+ // identities — in front of the one path that costs zero model calls. So this
277
+ // says what it saw and lets the run proceed, which also measures how often
278
+ // the mismatch is spurious. If it turns out to be rare, it can become a gate;
279
+ // deciding that by reasoning is how 174 got built in the first place.
280
+ let startedElsewhere = null;
281
+ if (flow.startScreen?.hash) {
282
+ try {
283
+ const here = await api.screenIdentity(device.udid, {});
284
+ if (here.hash && !graph.sameScreen(device.udid, here, flow.startScreen)) {
285
+ startedElsewhere = { recorded: flow.startScreen.hash.slice(0, 8), here: here.hash.slice(0, 8) };
286
+ }
287
+ } catch {
288
+ /* not knowing where we are is not a reason to refuse to try */
289
+ }
290
+ }
252
291
  const result = await runScript(device.udid, {
253
292
  steps: flow.steps,
254
293
  stopOnUnexpected: true,
@@ -256,11 +295,18 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
256
295
  minSteps: flow.minSteps ?? null,
257
296
  ...runOptions,
258
297
  });
259
- const ok = result.ranSteps === flow.steps.length;
298
+ // Same arithmetic as `saveFlow`, and the same defect: reaching the last step
299
+ // is not passing it. `result.ok` is the run's own verdict and was ignored
300
+ // here, so a replay that failed on its final assert promoted the flow it had
301
+ // just disproved. Both of a peer's saved flows were marked confirmed by
302
+ // replays that failed; the word "clean" in "confirmed by one clean replay"
303
+ // did not exist in the code path.
304
+ const everyStepPassed = (result.results ?? []).every((r) => r.ok);
305
+ const ok = result.ok !== false && everyStepPassed && result.ranSteps === flow.steps.length;
260
306
  // A clean replay is the confirmation a first traversal could not give.
261
307
  // Only when nothing was contradicted — a replay that ran to the end while
262
308
  // objecting to a step is not a promotion.
263
309
  const objected = (result.results ?? []).some((r) => CONTRADICTED.test(r.verification?.verdict ?? ''));
264
310
  const promoted = ok && !objected ? confirmFlow(device.udid, name) : false;
265
- return { ok, name, ...result, ...(promoted ? { promoted: true } : {}) };
311
+ return { ok, name, ...result, ...(promoted ? { promoted: true } : {}), ...(startedElsewhere ? { startedElsewhere } : {}) };
266
312
  }