simframe 0.16.0 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +42 -11
- package/package.json +1 -1
- package/scripts/ci-memory.mjs +24 -2
- package/src/actions.js +100 -2
- package/src/cli.js +17 -0
- package/src/graph.js +1 -1
- package/src/matching.js +3 -2
- package/src/mcp.js +24 -2
- package/src/navigate.js +49 -3
package/README.md
CHANGED
|
@@ -172,6 +172,18 @@ selector that did not *resolve* and nothing else, because retrying from a screen
|
|
|
172
172
|
you did not expect to be on is a second guess. A destructive-looking label is
|
|
173
173
|
refused as a substitute even if you list it.
|
|
174
174
|
|
|
175
|
+
**Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
|
|
176
|
+
"optional": true}` is skipped when nothing matches and runs normally when
|
|
177
|
+
something does. It exists because the opposite cost real time: a batch that
|
|
178
|
+
included a dismissal for a first-launch sheet lost **six correct steps** on the
|
|
179
|
+
next run because the sheet *did not appear*. A nag screen, a permission prompt,
|
|
180
|
+
a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
|
|
181
|
+
is the expensive outcome — a call per step instead of a call per flow. Only a
|
|
182
|
+
selector that resolved to **nothing** is absorbed: a target that is on screen
|
|
183
|
+
twice is ambiguous, not absent, and must still verify. Skipped steps are
|
|
184
|
+
reported, because "the sheet was gone" and "the sheet was dismissed" are
|
|
185
|
+
different facts.
|
|
186
|
+
|
|
175
187
|
**`{"seek": "change username", "budget": 6}`** opens containers, checks, and
|
|
176
188
|
comes back, depth first, inside a hard budget. It **acts** — opening a door
|
|
177
189
|
changes state — and it refuses to open anything that commits, abandons or
|
|
@@ -400,7 +412,7 @@ steer the model is a tool surface the model uses wrong.
|
|
|
400
412
|
| `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
|
|
401
413
|
| `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
|
|
402
414
|
| `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
|
|
403
|
-
| `sim_flow_run` | Replay a flow
|
|
415
|
+
| `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
|
|
404
416
|
| `sim_find` | Resolve an intent to one control, without acting on it. |
|
|
405
417
|
| `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
|
|
406
418
|
| `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
|
|
@@ -710,16 +722,35 @@ feels, and it took two field reports to see it. Per-step wall clock is
|
|
|
710
722
|
(model round trip + simframe work × n) / n for n steps in one call
|
|
711
723
|
```
|
|
712
724
|
|
|
713
|
-
|
|
714
|
-
|
|
715
|
-
|
|
716
|
-
|
|
717
|
-
|
|
|
718
|
-
|
|
|
719
|
-
| a
|
|
720
|
-
|
|
|
721
|
-
|
|
722
|
-
|
|
725
|
+
**Compare like with like, which this table used not to.** A human tester's
|
|
726
|
+
1.95 s per step is the *whole* loop — look at the screen, decide what to do,
|
|
727
|
+
do it. Any row that does not include a decision is not comparable to it.
|
|
728
|
+
|
|
729
|
+
| | perceives | decides | acts | per step |
|
|
730
|
+
| --- | --- | --- | --- | --- |
|
|
731
|
+
| **a human tester, measured** | yes | yes | yes | **1.95 s** |
|
|
732
|
+
| **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
|
|
733
|
+
| simframe, batch of 4 | yes | yes | yes | ~6.7 s |
|
|
734
|
+
| simframe, one model call per step | yes | yes | yes | ~21.7 s |
|
|
735
|
+
| — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
|
|
736
|
+
| — a saved flow replayed | yes | **no** | yes | 1.98 s |
|
|
737
|
+
|
|
738
|
+
**The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
|
|
739
|
+
a human**, and about eleven times when a failure forces one model call per step.
|
|
740
|
+
|
|
741
|
+
The two indented rows are the ones this README used to lead with, and both are
|
|
742
|
+
category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
|
|
743
|
+
taken out — the thinking is the model round trip, which is most of the clock.
|
|
744
|
+
And a replay *decides nothing*: it is a recording being played back, so its fair
|
|
745
|
+
counterpart is a human repeating a flow they have memorised, who would be well
|
|
746
|
+
under 1.95 s. Replay against a human working something out for the first time is
|
|
747
|
+
a rehearsal measured against a first attempt.
|
|
748
|
+
|
|
749
|
+
What the two rows do establish, and it is the finding that reordered this
|
|
750
|
+
project: **there is nothing left to win inside the engine.** 1.7 s is small
|
|
751
|
+
beside a 20 s round trip, so making perception or input faster buys single-digit
|
|
752
|
+
percentages. The only variable that matters is `n` — how many steps one decision
|
|
753
|
+
covers. Every improvement here has come from raising it, not from faster code.
|
|
723
754
|
A field report put the split at **34% simframe, 60% agent round trips** over
|
|
724
755
|
462 s of wall clock — the tester's *"30+ seconds between each step"* was
|
|
725
756
|
accurate and was not simframe. So there is nothing left to win inside the
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.17.0",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
5
|
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -646,12 +646,25 @@ console.log('\n--- what may be saved, and what may not ---');
|
|
|
646
646
|
// evidence for it. Checked against whatever the run actually produced rather
|
|
647
647
|
// than assuming a shape, as before.
|
|
648
648
|
const attempt = await jsonRetry(['do', LOOP, `--save=${FLOW_NAME}`], { allowFail: true });
|
|
649
|
+
//
|
|
650
|
+
// Three outcomes, not two. The missing one is a step that ran and *failed*: a
|
|
651
|
+
// failing step stops the batch, so a failure on the last step leaves
|
|
652
|
+
// `ranSteps === steps.length` and the run reads as complete. A peer watched
|
|
653
|
+
// `FLOW FAILED — 4 ok, 1 failed (of 5)` save itself. This branch is also the
|
|
654
|
+
// one that goes stale: yesterday the save contract was replaced, the unit test
|
|
655
|
+
// was updated, and this file was left asserting the old rule — so the three
|
|
656
|
+
// cases are enumerated here explicitly rather than left as an `else`.
|
|
649
657
|
const contradicted = attempt.results.some((r) => /^unexpected/.test(r.verification?.verdict ?? ''));
|
|
658
|
+
const anyFailed = attempt.results.some((r) => !r.ok);
|
|
650
659
|
const clean = attempt.results.every((r) => !r.verification || r.verification.verdict === 'ok');
|
|
651
660
|
if (contradicted) {
|
|
652
661
|
check(attempt.saved?.ok === false && attempt.saved?.reason === 'contradicted-steps',
|
|
653
662
|
'a flow with a contradicted step is refused, not quietly saved',
|
|
654
663
|
`${attempt.saved?.reason} (${(attempt.saved?.verdicts ?? []).join(', ')})`);
|
|
664
|
+
} else if (anyFailed) {
|
|
665
|
+
check(attempt.saved?.ok === false && attempt.saved?.reason === 'failed-steps',
|
|
666
|
+
'a flow with a step that ran and failed is refused, not quietly saved',
|
|
667
|
+
`${attempt.saved?.reason} (failed step ${(attempt.saved?.failed ?? []).join(', ')})`);
|
|
655
668
|
} else {
|
|
656
669
|
check(attempt.saved?.ok === true, 'a first traversal is recordable',
|
|
657
670
|
`${attempt.saved?.steps} steps`);
|
|
@@ -676,12 +689,21 @@ if (check(forced.saved?.ok === true, 'and --force saves it anyway', `${forced.sa
|
|
|
676
689
|
// state nothing ever leaves: a clean replay is the confirmation a first
|
|
677
690
|
// traversal could not give. Only asserted when the replay actually ran to
|
|
678
691
|
// the end — a partial replay promotes nothing, deliberately.
|
|
692
|
+
const after = await jsonRetry(['flow', 'list']);
|
|
693
|
+
const entry = after.find((f) => f.name === FLOW_NAME);
|
|
679
694
|
if (replayed.ok) {
|
|
680
|
-
const after = await jsonRetry(['flow', 'list']);
|
|
681
|
-
const entry = after.find((f) => f.name === FLOW_NAME);
|
|
682
695
|
check(entry && !entry.provisional,
|
|
683
696
|
'and a clean replay confirms a provisional flow',
|
|
684
697
|
`provisional=${entry?.provisional ?? 'gone'}`);
|
|
698
|
+
} else {
|
|
699
|
+
// The other direction, and it is the one that was silently wrong: promotion
|
|
700
|
+
// keyed on `ranSteps === steps.length`, which a failure on the *last* step
|
|
701
|
+
// satisfies. Both of a peer's saved flows were marked confirmed by replays
|
|
702
|
+
// that failed. A confirmation that a failing replay can grant is not a
|
|
703
|
+
// confirmation, so the negative case has to be checked too.
|
|
704
|
+
check(entry?.provisional === true,
|
|
705
|
+
'and a replay that failed does NOT confirm the flow it just disproved',
|
|
706
|
+
`ok=${replayed.ok}, provisional=${entry?.provisional}`);
|
|
685
707
|
}
|
|
686
708
|
const unknown = await cli(['flow', 'run', 'no-such-flow'], { expectFail: true });
|
|
687
709
|
check(/no flow/i.test(unknown), 'an unknown flow name is refused with what is known');
|
package/src/actions.js
CHANGED
|
@@ -346,6 +346,16 @@ export async function runScript(
|
|
|
346
346
|
// the caller returns to Claude is rendered from this, so describing the end
|
|
347
347
|
// state costs nothing beyond the verification pass the flow already ran.
|
|
348
348
|
let endScreen = null;
|
|
349
|
+
// Where this run began, so a flow saved from it can say what it assumes.
|
|
350
|
+
//
|
|
351
|
+
// Null when the first step does not depend on where we were — see
|
|
352
|
+
// `resetsTheScreen`.
|
|
353
|
+
//
|
|
354
|
+
// It was recorded as `null` on every flow 0.16.0 saved, which a peer caught:
|
|
355
|
+
// "replay never checks it is starting where it was recorded". Reading it is
|
|
356
|
+
// free — the first step computes the identity anyway — and it is a fact the
|
|
357
|
+
// recording had and discarded.
|
|
358
|
+
let startScreen = null;
|
|
349
359
|
// At most one recovery per run. Pressing home while already on the springboard
|
|
350
360
|
// moves nothing and is not a failure, so an unbounded retry would rebuild the
|
|
351
361
|
// session and press again on every such step for no reason.
|
|
@@ -393,6 +403,9 @@ export async function runScript(
|
|
|
393
403
|
? (carriedScreen ?? await api.screenIdentity(deviceQuery, { options, settleMs: stableMs, timeoutMs, confirmNovel }))
|
|
394
404
|
: null;
|
|
395
405
|
carriedScreen = null;
|
|
406
|
+
if (i === 0 && !resetsTheScreen(step)) {
|
|
407
|
+
startScreen = beforeScreen ? { hash: beforeScreen.hash ?? null, tokens: beforeScreen.tokens ?? null } : null;
|
|
408
|
+
}
|
|
396
409
|
// What this action did last time it was taken here, if ever.
|
|
397
410
|
const prediction = verify && beforeScreen?.hash ? graph.predict(udid, beforeScreen, step) : null;
|
|
398
411
|
try {
|
|
@@ -945,6 +958,33 @@ export async function runScript(
|
|
|
945
958
|
break;
|
|
946
959
|
}
|
|
947
960
|
} catch (err) {
|
|
961
|
+
// A step that was allowed not to be there, and is not there.
|
|
962
|
+
//
|
|
963
|
+
// `or:` supplies alternative *selectors*; nothing could express "skip
|
|
964
|
+
// this if absent". So a batch crossing a first-launch nag, a permission
|
|
965
|
+
// prompt or a "What's New" could not be batched at all — the interstitial
|
|
966
|
+
// appears on run one and not on run two, the step fails, and every
|
|
967
|
+
// remaining step is discarded. A peer lost six correct steps because a
|
|
968
|
+
// sheet *did not* appear, and on the framing that every hard-fail
|
|
969
|
+
// dropping a caller back to single-stepping is a latency bug, that is the
|
|
970
|
+
// most expensive shape there is: ~20 s per step, for the rest of the plan.
|
|
971
|
+
//
|
|
972
|
+
// Deliberately narrow. Only `unknown_screen` — the selector resolved to
|
|
973
|
+
// nothing — absorbs. `ambiguous_intent` means the target *is* there,
|
|
974
|
+
// twice, and a step that is present must still verify; so must one that
|
|
975
|
+
// resolved and then failed. "Skip if absent" must not become "tap
|
|
976
|
+
// whatever is there".
|
|
977
|
+
if (step.optional && didNotResolve(err)) {
|
|
978
|
+
results.push({
|
|
979
|
+
index: i,
|
|
980
|
+
action: step.action,
|
|
981
|
+
ok: true,
|
|
982
|
+
skipped: true,
|
|
983
|
+
ms: Date.now() - stepStart,
|
|
984
|
+
detail: `skipped — optional, and ${JSON.stringify(String(goalOf(step) ?? step.action))} is not on this screen`,
|
|
985
|
+
});
|
|
986
|
+
continue;
|
|
987
|
+
}
|
|
948
988
|
results.push({ index: i, action: step.action, ok: false, ms: Date.now() - stepStart, error: err.message });
|
|
949
989
|
const why = metrics.reasonForStepError(step, err);
|
|
950
990
|
noteEscalation({
|
|
@@ -1020,6 +1060,7 @@ export async function runScript(
|
|
|
1020
1060
|
// navigate.saveFlow without the caller reassembling what it just ran.
|
|
1021
1061
|
steps,
|
|
1022
1062
|
flowId,
|
|
1063
|
+
startScreen,
|
|
1023
1064
|
endScreen,
|
|
1024
1065
|
results,
|
|
1025
1066
|
ok: !failed,
|
|
@@ -1462,6 +1503,53 @@ export function alternativesFor(step) {
|
|
|
1462
1503
|
return list.map((v) => (typeof v === 'string' ? v : v?.value ?? v?.target ?? v?.label)).filter(Boolean);
|
|
1463
1504
|
}
|
|
1464
1505
|
|
|
1506
|
+
/**
|
|
1507
|
+
* Does this step make the screen we were on irrelevant?
|
|
1508
|
+
*
|
|
1509
|
+
* A flow that opens with `launch`, a deep link, or the home button does not
|
|
1510
|
+
* assume anything about where it starts — it goes and puts the device
|
|
1511
|
+
* somewhere. Recording a start screen for such a flow records where the device
|
|
1512
|
+
* happened to be beforehand, which is noise, and then reports a mismatch on
|
|
1513
|
+
* every replay that began anywhere else.
|
|
1514
|
+
*
|
|
1515
|
+
* Measured, not reasoned: a flow opening with `launch --relaunch` replayed
|
|
1516
|
+
* perfectly and still printed "recorded starting on 8292b488, replayed from
|
|
1517
|
+
* 39351dab", because the recording had followed a run that ended on Settings
|
|
1518
|
+
* root and the replay followed one that ended on About. Most flows open with a
|
|
1519
|
+
* launch, so that note would have fired on most correct replays — which is
|
|
1520
|
+
* item 175's defect exactly, in a feature added the same day 175 was written up.
|
|
1521
|
+
*/
|
|
1522
|
+
export function resetsTheScreen(step = {}) {
|
|
1523
|
+
if (step.action === 'launch' || step.action === 'openUrl') return true;
|
|
1524
|
+
return step.action === 'button' && /^home$/i.test(String(step.value ?? '').trim());
|
|
1525
|
+
}
|
|
1526
|
+
|
|
1527
|
+
/**
|
|
1528
|
+
* Did this failure mean "the thing is not here", as opposed to any other way a
|
|
1529
|
+
* step can fail? The one question `optional` is allowed to ask.
|
|
1530
|
+
*
|
|
1531
|
+
* **Not `reason === 'unknown_screen'`, which is what this said for one build and
|
|
1532
|
+
* what made `optional` a no-op on every screen it was built for.** The tag at
|
|
1533
|
+
* the throw site is chosen by whether the *screen* was recognised, not by
|
|
1534
|
+
* whether the *target* was found: on a screen recalled from memory an absent
|
|
1535
|
+
* target is tagged `ambiguous_intent` too. So a conditional interstitial on a
|
|
1536
|
+
* well-known screen — the whole use case — never absorbed. Caught by running it
|
|
1537
|
+
* against a device; the unit test exercised this predicate and passed, because
|
|
1538
|
+
* it tested the predicate rather than the path.
|
|
1539
|
+
*
|
|
1540
|
+
* The distinction that actually exists is `ambiguous`, set only where something
|
|
1541
|
+
* was found *several times over*, and the comment at that throw site had
|
|
1542
|
+
* already written down why it matters: "more time cannot make a thing unique,
|
|
1543
|
+
* and it can make an absent thing arrive." Absent is the case `optional` may
|
|
1544
|
+
* skip; present-twice must still verify.
|
|
1545
|
+
*/
|
|
1546
|
+
const RESOLVE_REASONS = new Set(['unknown_screen', 'ambiguous_intent']);
|
|
1547
|
+
|
|
1548
|
+
export function didNotResolve(err) {
|
|
1549
|
+
const tagged = metrics.escalationOf(err);
|
|
1550
|
+
return Boolean(tagged && RESOLVE_REASONS.has(tagged.reason) && !tagged.ambiguous);
|
|
1551
|
+
}
|
|
1552
|
+
|
|
1465
1553
|
export function mayRetryAfter(err) {
|
|
1466
1554
|
const tagged = metrics.escalationOf(err);
|
|
1467
1555
|
return Boolean(tagged && RESOLVE_FAILURES.has(tagged.reason));
|
|
@@ -3246,13 +3334,23 @@ export function settleEvidence(w) {
|
|
|
3246
3334
|
*/
|
|
3247
3335
|
export function flowSummary(res, { withTime = true } = {}) {
|
|
3248
3336
|
const time = withTime && Number.isFinite(res.totalMs) ? ` in ${res.totalMs}ms` : '';
|
|
3249
|
-
|
|
3337
|
+
// A step that was allowed not to be there and was not there did not run, and
|
|
3338
|
+
// saying so is the difference between "the nag screen was gone" and "the nag
|
|
3339
|
+
// screen was dismissed". Same lesson as the denominator above: a count that
|
|
3340
|
+
// quietly means two things is the defect.
|
|
3341
|
+
const skipped = (res.results ?? []).filter((r) => r.skipped).length;
|
|
3342
|
+
// After the time, in parentheses. Written inline before it, this read
|
|
3343
|
+
// "skipped as absent in 3412ms" — which says the skipping took 3.4 seconds.
|
|
3344
|
+
const skips = skipped
|
|
3345
|
+
? ` (${skipped} optional step${skipped === 1 ? '' : 's'} skipped as absent)`
|
|
3346
|
+
: '';
|
|
3347
|
+
if (res.ok) return `flow completed — ${res.ranSteps}/${res.totalSteps} steps${time}${skips}`;
|
|
3250
3348
|
const failed = (res.results ?? []).filter((r) => r.ok === false).length || 1;
|
|
3251
3349
|
const worked = Math.max(0, res.ranSteps - failed);
|
|
3252
3350
|
const unattempted = Math.max(0, res.totalSteps - res.ranSteps);
|
|
3253
3351
|
return `FLOW FAILED — ${worked} ok, ${failed} failed`
|
|
3254
3352
|
+ (unattempted ? `, ${unattempted} not attempted` : '')
|
|
3255
|
-
+ ` (of ${res.totalSteps})${time}`;
|
|
3353
|
+
+ ` (of ${res.totalSteps})${time}${skips}`;
|
|
3256
3354
|
}
|
|
3257
3355
|
|
|
3258
3356
|
/**
|
package/src/cli.js
CHANGED
|
@@ -285,6 +285,20 @@ async function main() {
|
|
|
285
285
|
// is a different machine and should say so rather than be guessed at.
|
|
286
286
|
if (flags.readyTimeoutMs) options.readyTimeoutMs = num(flags.readyTimeoutMs);
|
|
287
287
|
|
|
288
|
+
// `--help` on any command, not only as the command.
|
|
289
|
+
//
|
|
290
|
+
// `simframe hpi --help` did the two worst things in sequence: without
|
|
291
|
+
// `--device` it demanded a device *before* printing help, and with one it
|
|
292
|
+
// ignored `--help` and ran the report. A flag that is ignored is worse than
|
|
293
|
+
// one that is rejected — it silently does something other than what was
|
|
294
|
+
// asked. (The general case, unrecognised flags being dropped everywhere
|
|
295
|
+
// rather than here, is filed separately; this is the one that was measured
|
|
296
|
+
// costing a round trip.)
|
|
297
|
+
if (flags.help || flags.h) {
|
|
298
|
+
process.stdout.write(USAGE);
|
|
299
|
+
return;
|
|
300
|
+
}
|
|
301
|
+
|
|
288
302
|
switch (command) {
|
|
289
303
|
case undefined:
|
|
290
304
|
case '-h':
|
|
@@ -977,6 +991,9 @@ async function main() {
|
|
|
977
991
|
return;
|
|
978
992
|
}
|
|
979
993
|
emit(flags, res, [
|
|
994
|
+
...(res.startedElsewhere
|
|
995
|
+
? [`NOTE: recorded starting on ${res.startedElsewhere.recorded}, replayed from ${res.startedElsewhere.here}`]
|
|
996
|
+
: []),
|
|
980
997
|
...(res.results ?? []).map(stepLine),
|
|
981
998
|
actions.flowSummary(res, { withTime: false }),
|
|
982
999
|
]);
|
package/src/graph.js
CHANGED
|
@@ -869,7 +869,7 @@ export const VERDICTS = ['ok', 'no-visible-change', 'unexpected-screen', 'unveri
|
|
|
869
869
|
* that already exists do its job. A string still works and still means "exact
|
|
870
870
|
* match only", which is right for a stored prediction that has no tokens.
|
|
871
871
|
*/
|
|
872
|
-
function sameScreen(udid, a, b) {
|
|
872
|
+
export function sameScreen(udid, a, b) {
|
|
873
873
|
const hashOf = (v) => (typeof v === 'string' ? v : v?.hash);
|
|
874
874
|
const ha = hashOf(a);
|
|
875
875
|
const hb = hashOf(b);
|
package/src/matching.js
CHANGED
|
@@ -224,8 +224,9 @@ export function rank(targets, intent, { screen } = {}) {
|
|
|
224
224
|
|
|
225
225
|
// **A distinctive fragment of one long name, when nothing else came close.**
|
|
226
226
|
//
|
|
227
|
-
// Reported from the field: `waitFor
|
|
228
|
-
// whose own "Visible:" list printed
|
|
227
|
+
// Reported from the field: a `waitFor` on a seven-digit record number gave up
|
|
228
|
+
// after 20 s on a screen whose own "Visible:" list printed that number inside
|
|
229
|
+
// a heading — `Record #<digits>`, seventeen characters. The reporter's
|
|
229
230
|
// guess was that `#` was being treated as significant, or that the matcher
|
|
230
231
|
// was anchored. It is neither — the number scores **0.287** against the 0.45
|
|
231
232
|
// floor, because the substring branch in `nameScore` scales by how much of
|
package/src/mcp.js
CHANGED
|
@@ -134,7 +134,7 @@ const TOOLS = [
|
|
|
134
134
|
steps: {
|
|
135
135
|
type: 'array',
|
|
136
136
|
description:
|
|
137
|
-
'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
|
|
137
|
+
'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and "optional":true to any step that may legitimately have nothing to act on, which is how you cross a first-launch nag, a permission sheet or a "What\'s New" in one batch — {"tap":"Not Now","optional":true} is skipped when nothing matches and runs normally when it does, so the rest of the plan survives an interstitial that did not appear — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
|
|
138
138
|
items: { type: 'object' },
|
|
139
139
|
},
|
|
140
140
|
autoSettle: {
|
|
@@ -1077,11 +1077,25 @@ async function doScript(target, args, options) {
|
|
|
1077
1077
|
+ ' seen before to compare against. Replay it once with sim_flow_run and it'
|
|
1078
1078
|
+ ' is confirmed. A replay costs no model calls.'
|
|
1079
1079
|
: ' — confirmed; replay with sim_flow_run for zero model calls';
|
|
1080
|
+
// One sentence per reason, because "incomplete-run" and "failed-steps" ask
|
|
1081
|
+
// for opposite things from the caller and the bare reason names neither.
|
|
1082
|
+
const why = {
|
|
1083
|
+
'contradicted-steps': 'a step landed somewhere memory says it should not have,'
|
|
1084
|
+
+ ' so this route is not safe to replay. Check the flagged step before re-running it.',
|
|
1085
|
+
'failed-steps': 'a step ran and failed, so replaying this would fail the same way'
|
|
1086
|
+
+ ' every time. Fix the step and re-run the batch, then save.',
|
|
1087
|
+
'incomplete-run': 'the run stopped before its last step, so the rest of the route'
|
|
1088
|
+
+ ' was never observed. Re-run it to the end, then save.',
|
|
1089
|
+
}[saved.reason];
|
|
1090
|
+
// `verdicts` has a hole wherever a step never produced one — a failed step
|
|
1091
|
+
// has no verification — and joining those printed `(unverified, unverified, )`.
|
|
1092
|
+
const seen = (saved.verdicts ?? []).filter(Boolean);
|
|
1080
1093
|
lines.push(
|
|
1081
1094
|
saved.ok
|
|
1082
1095
|
? `saved as flow "${saved.name}" (${saved.steps} steps)${how}`
|
|
1083
1096
|
: `NOT saved as "${args.saveAs}": ${saved.reason}`
|
|
1084
|
-
+ `${
|
|
1097
|
+
+ `${why ? ` — ${why}` : ''}`
|
|
1098
|
+
+ `${seen.length ? ` (verdicts: ${seen.join(', ')})` : ''}`,
|
|
1085
1099
|
);
|
|
1086
1100
|
}
|
|
1087
1101
|
const escalated = (res.results ?? []).some((r) => metrics.ESCALATING_VERDICTS.has(r.verification?.verdict));
|
|
@@ -1167,6 +1181,14 @@ async function flowRun(target, args, options) {
|
|
|
1167
1181
|
};
|
|
1168
1182
|
}
|
|
1169
1183
|
const lines = [`flow "${args.name}"`, ...stepLines(res)];
|
|
1184
|
+
// Said out loud or not said at all. A flow now records the screen it was
|
|
1185
|
+
// recorded on, and a mismatch that only exists in the return value is a fact
|
|
1186
|
+
// nobody reads — which is the same shape as the check that could not fail.
|
|
1187
|
+
if (res.startedElsewhere) {
|
|
1188
|
+
lines.push(`NOTE: this flow was recorded starting on screen ${res.startedElsewhere.recorded},`
|
|
1189
|
+
+ ` and this replay started on ${res.startedElsewhere.here}. Not refused — content-driven`
|
|
1190
|
+
+ ' screens legitimately change identity — but if the steps below fail to resolve, this is why.');
|
|
1191
|
+
}
|
|
1170
1192
|
lines.push('', await mapFrom(target, options, res.endScreen, { verdictLine: verdictLineFor(res.results) }));
|
|
1171
1193
|
return { content: [text(lines.join('\n'))], isError: !res.ok };
|
|
1172
1194
|
}
|
package/src/navigate.js
CHANGED
|
@@ -170,14 +170,31 @@ const CONTRADICTED = /^unexpected/;
|
|
|
170
170
|
* and a flow that did not reach its own last step.
|
|
171
171
|
*/
|
|
172
172
|
export function saveFlow(udid, name, script, { force = false } = {}) {
|
|
173
|
-
const
|
|
173
|
+
const results = script.results ?? [];
|
|
174
|
+
const verdicts = results.map((r) => r.verification?.verdict);
|
|
174
175
|
const contradicted = verdicts.filter((v) => v && CONTRADICTED.test(v));
|
|
175
176
|
const ranAll = script.ranSteps == null
|
|
176
177
|
|| script.steps == null
|
|
177
178
|
|| script.ranSteps >= (script.steps?.length ?? 0);
|
|
179
|
+
// A step that ran and failed is not a step that ran.
|
|
180
|
+
//
|
|
181
|
+
// `ranSteps` counts steps *attempted*, and a failing step stops the batch —
|
|
182
|
+
// so a run whose failure is on the **last** step has `ranSteps === steps.length`
|
|
183
|
+
// and passed the gate above. A peer watched `FLOW FAILED — 4 ok, 1 failed (of
|
|
184
|
+
// 5)` save itself as a flow, which is a flow guaranteed to fail forever. The
|
|
185
|
+
// predicate wanted "every step succeeded" and was written as "every step was
|
|
186
|
+
// reached"; on every run except one they are the same sentence.
|
|
187
|
+
//
|
|
188
|
+
// `!r.ok` rather than `r.ok === false` on purpose: a result that does not say
|
|
189
|
+
// it succeeded has not said it succeeded, and a gate that only catches an
|
|
190
|
+
// explicit `false` is one absent field away from being unable to fail.
|
|
191
|
+
const failedSteps = results.filter((r) => !r.ok);
|
|
178
192
|
if (!force && contradicted.length) {
|
|
179
193
|
return { ok: false, reason: 'contradicted-steps', verdicts };
|
|
180
194
|
}
|
|
195
|
+
if (!force && failedSteps.length) {
|
|
196
|
+
return { ok: false, reason: 'failed-steps', verdicts, failed: failedSteps.map((r) => r.index ?? null) };
|
|
197
|
+
}
|
|
181
198
|
if (!force && !ranAll) {
|
|
182
199
|
return { ok: false, reason: 'incomplete-run', verdicts };
|
|
183
200
|
}
|
|
@@ -249,6 +266,28 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
|
|
|
249
266
|
// a human doing the same thing. `minSteps` comes from the flow definition or
|
|
250
267
|
// stays null — the step count of a recorded route is not a claim about the
|
|
251
268
|
// shortest one.
|
|
269
|
+
// Is this the screen the flow was recorded on?
|
|
270
|
+
//
|
|
271
|
+
// Reported, not refused, and the distinction is deliberate. A replay from the
|
|
272
|
+
// wrong screen is mostly self-limiting — the first selector does not resolve
|
|
273
|
+
// and the run stops one step in, and the destructive vocabulary is barred by
|
|
274
|
+
// the verify barrier either way. Refusing on a mismatch would put the exact
|
|
275
|
+
// failure mode of item 174 — content-driven screens fragmenting into several
|
|
276
|
+
// identities — in front of the one path that costs zero model calls. So this
|
|
277
|
+
// says what it saw and lets the run proceed, which also measures how often
|
|
278
|
+
// the mismatch is spurious. If it turns out to be rare, it can become a gate;
|
|
279
|
+
// deciding that by reasoning is how 174 got built in the first place.
|
|
280
|
+
let startedElsewhere = null;
|
|
281
|
+
if (flow.startScreen?.hash) {
|
|
282
|
+
try {
|
|
283
|
+
const here = await api.screenIdentity(device.udid, {});
|
|
284
|
+
if (here.hash && !graph.sameScreen(device.udid, here, flow.startScreen)) {
|
|
285
|
+
startedElsewhere = { recorded: flow.startScreen.hash.slice(0, 8), here: here.hash.slice(0, 8) };
|
|
286
|
+
}
|
|
287
|
+
} catch {
|
|
288
|
+
/* not knowing where we are is not a reason to refuse to try */
|
|
289
|
+
}
|
|
290
|
+
}
|
|
252
291
|
const result = await runScript(device.udid, {
|
|
253
292
|
steps: flow.steps,
|
|
254
293
|
stopOnUnexpected: true,
|
|
@@ -256,11 +295,18 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
|
|
|
256
295
|
minSteps: flow.minSteps ?? null,
|
|
257
296
|
...runOptions,
|
|
258
297
|
});
|
|
259
|
-
|
|
298
|
+
// Same arithmetic as `saveFlow`, and the same defect: reaching the last step
|
|
299
|
+
// is not passing it. `result.ok` is the run's own verdict and was ignored
|
|
300
|
+
// here, so a replay that failed on its final assert promoted the flow it had
|
|
301
|
+
// just disproved. Both of a peer's saved flows were marked confirmed by
|
|
302
|
+
// replays that failed; the word "clean" in "confirmed by one clean replay"
|
|
303
|
+
// did not exist in the code path.
|
|
304
|
+
const everyStepPassed = (result.results ?? []).every((r) => r.ok);
|
|
305
|
+
const ok = result.ok !== false && everyStepPassed && result.ranSteps === flow.steps.length;
|
|
260
306
|
// A clean replay is the confirmation a first traversal could not give.
|
|
261
307
|
// Only when nothing was contradicted — a replay that ran to the end while
|
|
262
308
|
// objecting to a step is not a promotion.
|
|
263
309
|
const objected = (result.results ?? []).some((r) => CONTRADICTED.test(r.verification?.verdict ?? ''));
|
|
264
310
|
const promoted = ok && !objected ? confirmFlow(device.udid, name) : false;
|
|
265
|
-
return { ok, name, ...result, ...(promoted ? { promoted: true } : {}) };
|
|
311
|
+
return { ok, name, ...result, ...(promoted ? { promoted: true } : {}), ...(startedElsewhere ? { startedElsewhere } : {}) };
|
|
266
312
|
}
|