simframe 0.15.1 → 0.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +62 -1
- package/package.json +1 -1
- package/scripts/bench-hpi.mjs +6 -3
- package/scripts/ci-memory.mjs +55 -10
- package/src/actions.js +245 -15
- package/src/cli.js +25 -0
- package/src/graph.js +1 -1
- package/src/matching.js +43 -1
- package/src/mcp.js +55 -3
- package/src/metrics.js +81 -6
- package/src/navigate.js +125 -9
package/README.md
CHANGED
|
@@ -172,6 +172,18 @@ selector that did not *resolve* and nothing else, because retrying from a screen
|
|
|
172
172
|
you did not expect to be on is a second guess. A destructive-looking label is
|
|
173
173
|
refused as a substitute even if you list it.
|
|
174
174
|
|
|
175
|
+
**Steps that may legitimately have nothing to act on.** `{"tap": "Not Now",
|
|
176
|
+
"optional": true}` is skipped when nothing matches and runs normally when
|
|
177
|
+
something does. It exists because the opposite cost real time: a batch that
|
|
178
|
+
included a dismissal for a first-launch sheet lost **six correct steps** on the
|
|
179
|
+
next run because the sheet *did not appear*. A nag screen, a permission prompt,
|
|
180
|
+
a "What's New" or a cold-start splash otherwise makes a flow unbatchable, which
|
|
181
|
+
is the expensive outcome — a call per step instead of a call per flow. Only a
|
|
182
|
+
selector that resolved to **nothing** is absorbed: a target that is on screen
|
|
183
|
+
twice is ambiguous, not absent, and must still verify. Skipped steps are
|
|
184
|
+
reported, because "the sheet was gone" and "the sheet was dismissed" are
|
|
185
|
+
different facts.
|
|
186
|
+
|
|
175
187
|
**`{"seek": "change username", "budget": 6}`** opens containers, checks, and
|
|
176
188
|
comes back, depth first, inside a hard budget. It **acts** — opening a door
|
|
177
189
|
changes state — and it refuses to open anything that commits, abandons or
|
|
@@ -400,7 +412,7 @@ steer the model is a tool surface the model uses wrong.
|
|
|
400
412
|
| `sim_do` | **The main tool.** A whole flow in one call — tap, type, scroll, wait, assert — each step settling before the next and verified against what it did last time. |
|
|
401
413
|
| `sim_state` | The cheapest question there is: has anything changed **since your last look**, and which regions moved. |
|
|
402
414
|
| `sim_goto` | Walk to a screen simframe has been to before, planning the route through remembered transitions. |
|
|
403
|
-
| `sim_flow_run` | Replay a flow
|
|
415
|
+
| `sim_flow_run` | Replay a saved flow — **zero model calls**, which is the only path to human wall clock. A first traversal saves as *provisional*; one replay in which every step passed confirms it. A run with a contradicted step, a failed step, or one that never reached its last step is refused and says which. |
|
|
404
416
|
| `sim_find` | Resolve an intent to one control, without acting on it. |
|
|
405
417
|
| `sim_tap` · `sim_type_into` · `sim_scroll_to` · `sim_wait_for` · `sim_assert` | Single actions, for when you genuinely only have one step. Each is one `sim_do` step underneath. |
|
|
406
418
|
| `sim_launch` · `sim_open_url` · `sim_permission` | Launch with arguments and environment; open a deep link; grant a privacy permission instead of tapping a system alert. A launch is **confirmed to have reached the front**, by comparing the pid `simctl` started against the pid the device reports as frontmost — so *"the process started"* is no longer reported as *"the app is on screen"*. |
|
|
@@ -701,6 +713,55 @@ Reproduce all of it with `npm run bench`, which prints the same table against
|
|
|
701
713
|
your machine. Full detail, including the measurement traps, is in
|
|
702
714
|
[`docs/BENCHMARKS.md`](docs/BENCHMARKS.md).
|
|
703
715
|
|
|
716
|
+
### Wall clock per step — where the time really goes
|
|
717
|
+
|
|
718
|
+
Every number above is microscopic next to the one that decides how fast this
|
|
719
|
+
feels, and it took two field reports to see it. Per-step wall clock is
|
|
720
|
+
|
|
721
|
+
```
|
|
722
|
+
(model round trip + simframe work × n) / n for n steps in one call
|
|
723
|
+
```
|
|
724
|
+
|
|
725
|
+
**Compare like with like, which this table used not to.** A human tester's
|
|
726
|
+
1.95 s per step is the *whole* loop — look at the screen, decide what to do,
|
|
727
|
+
do it. Any row that does not include a decision is not comparable to it.
|
|
728
|
+
|
|
729
|
+
| | perceives | decides | acts | per step |
|
|
730
|
+
| --- | --- | --- | --- | --- |
|
|
731
|
+
| **a human tester, measured** | yes | yes | yes | **1.95 s** |
|
|
732
|
+
| **simframe, batch of 2 — the recorded median** | yes | yes | yes | **~11.7 s** |
|
|
733
|
+
| simframe, batch of 4 | yes | yes | yes | ~6.7 s |
|
|
734
|
+
| simframe, one model call per step | yes | yes | yes | ~21.7 s |
|
|
735
|
+
| — simframe's mechanical half alone | yes | **no** | yes | ~1.7 s |
|
|
736
|
+
| — a saved flow replayed | yes | **no** | yes | 1.98 s |
|
|
737
|
+
|
|
738
|
+
**The only like-for-like comparison is 1.95 s against ~11.7 s: about six times
|
|
739
|
+
a human**, and about eleven times when a failure forces one model call per step.
|
|
740
|
+
|
|
741
|
+
The two indented rows are the ones this README used to lead with, and both are
|
|
742
|
+
category errors when set against 1.95 s. The 1.7 s is simframe with the thinking
|
|
743
|
+
taken out — the thinking is the model round trip, which is most of the clock.
|
|
744
|
+
And a replay *decides nothing*: it is a recording being played back, so its fair
|
|
745
|
+
counterpart is a human repeating a flow they have memorised, who would be well
|
|
746
|
+
under 1.95 s. Replay against a human working something out for the first time is
|
|
747
|
+
a rehearsal measured against a first attempt.
|
|
748
|
+
|
|
749
|
+
What the two rows do establish, and it is the finding that reordered this
|
|
750
|
+
project: **there is nothing left to win inside the engine.** 1.7 s is small
|
|
751
|
+
beside a 20 s round trip, so making perception or input faster buys single-digit
|
|
752
|
+
percentages. The only variable that matters is `n` — how many steps one decision
|
|
753
|
+
covers. Every improvement here has come from raising it, not from faster code.
|
|
754
|
+
A field report put the split at **34% simframe, 60% agent round trips** over
|
|
755
|
+
462 s of wall clock — the tester's *"30+ seconds between each step"* was
|
|
756
|
+
accurate and was not simframe. So there is nothing left to win inside the
|
|
757
|
+
engine, and the only variable is `n`: `simframe hpi` reports `steps_per_call`
|
|
758
|
+
for exactly that reason.
|
|
759
|
+
|
|
760
|
+
Which is why **every hard-fail that drops a caller back to single-stepping is a
|
|
761
|
+
latency bug**. Between two field reports on the same flow, one run took 33 tool
|
|
762
|
+
calls and the next took **16**, for 28 executed steps and no screenshots at
|
|
763
|
+
all — the difference being defects fixed, not code made faster.
|
|
764
|
+
|
|
704
765
|
## How it works
|
|
705
766
|
|
|
706
767
|
```
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "simframe",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.17.0",
|
|
4
4
|
"mcpName": "io.github.lvlrSajjad/simframe",
|
|
5
5
|
"description": "Always-warm iOS Simulator and Android emulator frames: agents read the screen in ~20ms instead of waiting on screenshots. MCP server + CLI.",
|
|
6
6
|
"keywords": [
|
package/scripts/bench-hpi.mjs
CHANGED
|
@@ -243,8 +243,11 @@ const base = committed.overall ?? {};
|
|
|
243
243
|
const failures = metrics.gateAgainst(committed, report);
|
|
244
244
|
|
|
245
245
|
console.log(`\ngate vs ${path.relative(ROOT, baselineFile)} (measured ${committed.measured_at ?? '?'})`);
|
|
246
|
-
console.log(` HPI_accuracy ${base.hpi_accuracy ?? '—'} -> ${o.hpi_accuracy ?? '—'}`);
|
|
247
|
-
console.log(`
|
|
248
|
-
|
|
246
|
+
console.log(` HPI_accuracy ${base.hpi_accuracy ?? '—'} -> ${o.hpi_accuracy ?? '—'} (any drop fails)`);
|
|
247
|
+
console.log(` step_ratio ${o.step_ratio ?? '—'} (fails above ${metrics.STEP_RATIO_CEILING})`);
|
|
248
|
+
// Printed apart from the two that gate, and labelled, because a number in a
|
|
249
|
+
// gate block gets read as a threshold whether or not it is one.
|
|
250
|
+
console.log(` ${metrics.timeTrend(committed, report)}`);
|
|
251
|
+
console.log(` median of ${passTimes.length || 1} pass(es), on this host`);
|
|
249
252
|
for (const f of failures) console.log(`FAIL ${f}`);
|
|
250
253
|
process.exit(failures.length ? 1 : 0);
|
package/scripts/ci-memory.mjs
CHANGED
|
@@ -635,19 +635,44 @@ check(walked.ok === true || outcomes.includes(walked.reason),
|
|
|
635
635
|
walked.ok ? (walked.already ? 'already there' : `walked ${walked.ranSteps} step(s)`) : walked.reason);
|
|
636
636
|
|
|
637
637
|
console.log('\n--- what may be saved, and what may not ---');
|
|
638
|
-
// A
|
|
639
|
-
//
|
|
640
|
-
//
|
|
641
|
-
//
|
|
638
|
+
// **A first traversal must be recordable**, which is the opposite of what this
|
|
639
|
+
// check used to assert. The old contract refused any flow with a non-`ok`
|
|
640
|
+
// verdict — and a first traversal is all-`unverified` by construction, since
|
|
641
|
+
// there is no prior observation to compare against. So no flow could ever be
|
|
642
|
+
// recorded, replay was unreachable, and the only zero-model-call path in the
|
|
643
|
+
// tool was sealed shut behind a gate nobody could pass.
|
|
644
|
+
//
|
|
645
|
+
// The refusal now needs evidence *against* a step rather than the absence of
|
|
646
|
+
// evidence for it. Checked against whatever the run actually produced rather
|
|
647
|
+
// than assuming a shape, as before.
|
|
642
648
|
const attempt = await jsonRetry(['do', LOOP, `--save=${FLOW_NAME}`], { allowFail: true });
|
|
649
|
+
//
|
|
650
|
+
// Three outcomes, not two. The missing one is a step that ran and *failed*: a
|
|
651
|
+
// failing step stops the batch, so a failure on the last step leaves
|
|
652
|
+
// `ranSteps === steps.length` and the run reads as complete. A peer watched
|
|
653
|
+
// `FLOW FAILED — 4 ok, 1 failed (of 5)` save itself. This branch is also the
|
|
654
|
+
// one that goes stale: yesterday the save contract was replaced, the unit test
|
|
655
|
+
// was updated, and this file was left asserting the old rule — so the three
|
|
656
|
+
// cases are enumerated here explicitly rather than left as an `else`.
|
|
657
|
+
const contradicted = attempt.results.some((r) => /^unexpected/.test(r.verification?.verdict ?? ''));
|
|
658
|
+
const anyFailed = attempt.results.some((r) => !r.ok);
|
|
643
659
|
const clean = attempt.results.every((r) => !r.verification || r.verification.verdict === 'ok');
|
|
644
|
-
if (
|
|
645
|
-
check(attempt.saved?.ok ===
|
|
646
|
-
|
|
647
|
-
} else {
|
|
648
|
-
check(attempt.saved?.ok === false && attempt.saved?.reason === 'unverified-steps',
|
|
649
|
-
'a flow with an unverified step is refused, not quietly saved',
|
|
660
|
+
if (contradicted) {
|
|
661
|
+
check(attempt.saved?.ok === false && attempt.saved?.reason === 'contradicted-steps',
|
|
662
|
+
'a flow with a contradicted step is refused, not quietly saved',
|
|
650
663
|
`${attempt.saved?.reason} (${(attempt.saved?.verdicts ?? []).join(', ')})`);
|
|
664
|
+
} else if (anyFailed) {
|
|
665
|
+
check(attempt.saved?.ok === false && attempt.saved?.reason === 'failed-steps',
|
|
666
|
+
'a flow with a step that ran and failed is refused, not quietly saved',
|
|
667
|
+
`${attempt.saved?.reason} (failed step ${(attempt.saved?.failed ?? []).join(', ')})`);
|
|
668
|
+
} else {
|
|
669
|
+
check(attempt.saved?.ok === true, 'a first traversal is recordable',
|
|
670
|
+
`${attempt.saved?.steps} steps`);
|
|
671
|
+
check(attempt.saved?.provisional === !clean,
|
|
672
|
+
clean
|
|
673
|
+
? 'and a flow whose every step verified is confirmed outright'
|
|
674
|
+
: 'and it is marked provisional, because nothing had been seen before to compare against',
|
|
675
|
+
`provisional=${attempt.saved?.provisional}`);
|
|
651
676
|
}
|
|
652
677
|
|
|
653
678
|
// --force is the deliberate override, and it is what lets the replay machinery
|
|
@@ -660,6 +685,26 @@ if (check(forced.saved?.ok === true, 'and --force saves it anyway', `${forced.sa
|
|
|
660
685
|
check(replayed.ranSteps >= 1 && Array.isArray(replayed.results),
|
|
661
686
|
'and replays from disk with no model in the loop',
|
|
662
687
|
`${replayed.ranSteps}/${replayed.totalSteps} steps`);
|
|
688
|
+
// The other half of the bootstrap, and the reason "provisional" is not a
|
|
689
|
+
// state nothing ever leaves: a clean replay is the confirmation a first
|
|
690
|
+
// traversal could not give. Only asserted when the replay actually ran to
|
|
691
|
+
// the end — a partial replay promotes nothing, deliberately.
|
|
692
|
+
const after = await jsonRetry(['flow', 'list']);
|
|
693
|
+
const entry = after.find((f) => f.name === FLOW_NAME);
|
|
694
|
+
if (replayed.ok) {
|
|
695
|
+
check(entry && !entry.provisional,
|
|
696
|
+
'and a clean replay confirms a provisional flow',
|
|
697
|
+
`provisional=${entry?.provisional ?? 'gone'}`);
|
|
698
|
+
} else {
|
|
699
|
+
// The other direction, and it is the one that was silently wrong: promotion
|
|
700
|
+
// keyed on `ranSteps === steps.length`, which a failure on the *last* step
|
|
701
|
+
// satisfies. Both of a peer's saved flows were marked confirmed by replays
|
|
702
|
+
// that failed. A confirmation that a failing replay can grant is not a
|
|
703
|
+
// confirmation, so the negative case has to be checked too.
|
|
704
|
+
check(entry?.provisional === true,
|
|
705
|
+
'and a replay that failed does NOT confirm the flow it just disproved',
|
|
706
|
+
`ok=${replayed.ok}, provisional=${entry?.provisional}`);
|
|
707
|
+
}
|
|
663
708
|
const unknown = await cli(['flow', 'run', 'no-such-flow'], { expectFail: true });
|
|
664
709
|
check(/no flow/i.test(unknown), 'an unknown flow name is refused with what is known');
|
|
665
710
|
}
|
package/src/actions.js
CHANGED
|
@@ -19,6 +19,29 @@ import { launchApp, openUrl, setPermission, terminateApp } from './platform/inde
|
|
|
19
19
|
|
|
20
20
|
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
|
|
21
21
|
const MAX_PAUSE_MS = 5000;
|
|
22
|
+
/**
|
|
23
|
+
* Above this many characters a fill uses the pasteboard; at or below it, the
|
|
24
|
+
* keyboard.
|
|
25
|
+
*
|
|
26
|
+
* The pasteboard is the right default for long text and was the default for
|
|
27
|
+
* *all* text in `sweep`, which a field report showed costing far more than it
|
|
28
|
+
* saved. iOS 26 raises a system consent alert — *"… would like to paste from
|
|
29
|
+
* CoreSimulatorBridge"* — on the first paste into an app, and it fired on a
|
|
30
|
+
* **nine-character** value. The alert covered the form, collapsed the
|
|
31
|
+
* accessibility tree to OCR-only, and failed the batch; the recovery chain that
|
|
32
|
+
* followed was close to half of that session's wasted model round trips.
|
|
33
|
+
*
|
|
34
|
+
* `sim_type_into` already defaults to the keyboard. Only `sweep`'s fill did
|
|
35
|
+
* not, so the reporter got the pasteboard without asking for it.
|
|
36
|
+
*
|
|
37
|
+
* Why a threshold is safe rather than a trade against exactness — the reason
|
|
38
|
+
* pasteboard-by-default was chosen: `type` with `into` reads the field back,
|
|
39
|
+
* and retries once locally if nothing landed. So a keystroke path that a
|
|
40
|
+
* keyboard layout mangles is *caught*, not silently accepted. 40 characters is
|
|
41
|
+
* the reporter's own suggestion and covers names, emails and short
|
|
42
|
+
* descriptions, while a paragraph still goes by pasteboard where it belongs.
|
|
43
|
+
*/
|
|
44
|
+
const KEYBOARD_UP_TO = 40;
|
|
22
45
|
/**
|
|
23
46
|
* A tapped field is typed into once the screen has settled, not after a fixed
|
|
24
47
|
* wait.
|
|
@@ -323,6 +346,16 @@ export async function runScript(
|
|
|
323
346
|
// the caller returns to Claude is rendered from this, so describing the end
|
|
324
347
|
// state costs nothing beyond the verification pass the flow already ran.
|
|
325
348
|
let endScreen = null;
|
|
349
|
+
// Where this run began, so a flow saved from it can say what it assumes.
|
|
350
|
+
//
|
|
351
|
+
// Null when the first step does not depend on where we were — see
|
|
352
|
+
// `resetsTheScreen`.
|
|
353
|
+
//
|
|
354
|
+
// It was recorded as `null` on every flow 0.16.0 saved, which a peer caught:
|
|
355
|
+
// "replay never checks it is starting where it was recorded". Reading it is
|
|
356
|
+
// free — the first step computes the identity anyway — and it is a fact the
|
|
357
|
+
// recording had and discarded.
|
|
358
|
+
let startScreen = null;
|
|
326
359
|
// At most one recovery per run. Pressing home while already on the springboard
|
|
327
360
|
// moves nothing and is not a failure, so an unbounded retry would rebuild the
|
|
328
361
|
// session and press again on every such step for no reason.
|
|
@@ -370,6 +403,9 @@ export async function runScript(
|
|
|
370
403
|
? (carriedScreen ?? await api.screenIdentity(deviceQuery, { options, settleMs: stableMs, timeoutMs, confirmNovel }))
|
|
371
404
|
: null;
|
|
372
405
|
carriedScreen = null;
|
|
406
|
+
if (i === 0 && !resetsTheScreen(step)) {
|
|
407
|
+
startScreen = beforeScreen ? { hash: beforeScreen.hash ?? null, tokens: beforeScreen.tokens ?? null } : null;
|
|
408
|
+
}
|
|
373
409
|
// What this action did last time it was taken here, if ever.
|
|
374
410
|
const prediction = verify && beforeScreen?.hash ? graph.predict(udid, beforeScreen, step) : null;
|
|
375
411
|
try {
|
|
@@ -922,6 +958,33 @@ export async function runScript(
|
|
|
922
958
|
break;
|
|
923
959
|
}
|
|
924
960
|
} catch (err) {
|
|
961
|
+
// A step that was allowed not to be there, and is not there.
|
|
962
|
+
//
|
|
963
|
+
// `or:` supplies alternative *selectors*; nothing could express "skip
|
|
964
|
+
// this if absent". So a batch crossing a first-launch nag, a permission
|
|
965
|
+
// prompt or a "What's New" could not be batched at all — the interstitial
|
|
966
|
+
// appears on run one and not on run two, the step fails, and every
|
|
967
|
+
// remaining step is discarded. A peer lost six correct steps because a
|
|
968
|
+
// sheet *did not* appear, and on the framing that every hard-fail
|
|
969
|
+
// dropping a caller back to single-stepping is a latency bug, that is the
|
|
970
|
+
// most expensive shape there is: ~20 s per step, for the rest of the plan.
|
|
971
|
+
//
|
|
972
|
+
// Deliberately narrow. Only `unknown_screen` — the selector resolved to
|
|
973
|
+
// nothing — absorbs. `ambiguous_intent` means the target *is* there,
|
|
974
|
+
// twice, and a step that is present must still verify; so must one that
|
|
975
|
+
// resolved and then failed. "Skip if absent" must not become "tap
|
|
976
|
+
// whatever is there".
|
|
977
|
+
if (step.optional && didNotResolve(err)) {
|
|
978
|
+
results.push({
|
|
979
|
+
index: i,
|
|
980
|
+
action: step.action,
|
|
981
|
+
ok: true,
|
|
982
|
+
skipped: true,
|
|
983
|
+
ms: Date.now() - stepStart,
|
|
984
|
+
detail: `skipped — optional, and ${JSON.stringify(String(goalOf(step) ?? step.action))} is not on this screen`,
|
|
985
|
+
});
|
|
986
|
+
continue;
|
|
987
|
+
}
|
|
925
988
|
results.push({ index: i, action: step.action, ok: false, ms: Date.now() - stepStart, error: err.message });
|
|
926
989
|
const why = metrics.reasonForStepError(step, err);
|
|
927
990
|
noteEscalation({
|
|
@@ -997,6 +1060,7 @@ export async function runScript(
|
|
|
997
1060
|
// navigate.saveFlow without the caller reassembling what it just ran.
|
|
998
1061
|
steps,
|
|
999
1062
|
flowId,
|
|
1063
|
+
startScreen,
|
|
1000
1064
|
endScreen,
|
|
1001
1065
|
results,
|
|
1002
1066
|
ok: !failed,
|
|
@@ -1439,6 +1503,53 @@ export function alternativesFor(step) {
|
|
|
1439
1503
|
return list.map((v) => (typeof v === 'string' ? v : v?.value ?? v?.target ?? v?.label)).filter(Boolean);
|
|
1440
1504
|
}
|
|
1441
1505
|
|
|
1506
|
+
/**
|
|
1507
|
+
* Does this step make the screen we were on irrelevant?
|
|
1508
|
+
*
|
|
1509
|
+
* A flow that opens with `launch`, a deep link, or the home button does not
|
|
1510
|
+
* assume anything about where it starts — it goes and puts the device
|
|
1511
|
+
* somewhere. Recording a start screen for such a flow records where the device
|
|
1512
|
+
* happened to be beforehand, which is noise, and then reports a mismatch on
|
|
1513
|
+
* every replay that began anywhere else.
|
|
1514
|
+
*
|
|
1515
|
+
* Measured, not reasoned: a flow opening with `launch --relaunch` replayed
|
|
1516
|
+
* perfectly and still printed "recorded starting on 8292b488, replayed from
|
|
1517
|
+
* 39351dab", because the recording had followed a run that ended on Settings
|
|
1518
|
+
* root and the replay followed one that ended on About. Most flows open with a
|
|
1519
|
+
* launch, so that note would have fired on most correct replays — which is
|
|
1520
|
+
* item 175's defect exactly, in a feature added the same day 175 was written up.
|
|
1521
|
+
*/
|
|
1522
|
+
export function resetsTheScreen(step = {}) {
|
|
1523
|
+
if (step.action === 'launch' || step.action === 'openUrl') return true;
|
|
1524
|
+
return step.action === 'button' && /^home$/i.test(String(step.value ?? '').trim());
|
|
1525
|
+
}
|
|
1526
|
+
|
|
1527
|
+
/**
|
|
1528
|
+
* Did this failure mean "the thing is not here", as opposed to any other way a
|
|
1529
|
+
* step can fail? The one question `optional` is allowed to ask.
|
|
1530
|
+
*
|
|
1531
|
+
* **Not `reason === 'unknown_screen'`, which is what this said for one build and
|
|
1532
|
+
* what made `optional` a no-op on every screen it was built for.** The tag at
|
|
1533
|
+
* the throw site is chosen by whether the *screen* was recognised, not by
|
|
1534
|
+
* whether the *target* was found: on a screen recalled from memory an absent
|
|
1535
|
+
* target is tagged `ambiguous_intent` too. So a conditional interstitial on a
|
|
1536
|
+
* well-known screen — the whole use case — never absorbed. Caught by running it
|
|
1537
|
+
* against a device; the unit test exercised this predicate and passed, because
|
|
1538
|
+
* it tested the predicate rather than the path.
|
|
1539
|
+
*
|
|
1540
|
+
* The distinction that actually exists is `ambiguous`, set only where something
|
|
1541
|
+
* was found *several times over*, and the comment at that throw site had
|
|
1542
|
+
* already written down why it matters: "more time cannot make a thing unique,
|
|
1543
|
+
* and it can make an absent thing arrive." Absent is the case `optional` may
|
|
1544
|
+
* skip; present-twice must still verify.
|
|
1545
|
+
*/
|
|
1546
|
+
const RESOLVE_REASONS = new Set(['unknown_screen', 'ambiguous_intent']);
|
|
1547
|
+
|
|
1548
|
+
export function didNotResolve(err) {
|
|
1549
|
+
const tagged = metrics.escalationOf(err);
|
|
1550
|
+
return Boolean(tagged && RESOLVE_REASONS.has(tagged.reason) && !tagged.ambiguous);
|
|
1551
|
+
}
|
|
1552
|
+
|
|
1442
1553
|
export function mayRetryAfter(err) {
|
|
1443
1554
|
const tagged = metrics.escalationOf(err);
|
|
1444
1555
|
return Boolean(tagged && RESOLVE_FAILURES.has(tagged.reason));
|
|
@@ -1559,12 +1670,35 @@ async function stillOnPlan(deviceQuery, verification, nextStep, options) {
|
|
|
1559
1670
|
// neither is evidence about where we are.
|
|
1560
1671
|
if (!target || /^#\d+$/.test(String(target).trim()) || /^@?-?\d+\s*,\s*-?\d+$/.test(String(target).trim())) return false;
|
|
1561
1672
|
if (!ACTION_STEPS.has(nextStep.action) && nextStep.action !== 'assert' && nextStep.action !== 'waitFor') return false;
|
|
1562
|
-
|
|
1563
|
-
|
|
1564
|
-
|
|
1565
|
-
|
|
1566
|
-
|
|
1673
|
+
// **Settle before looking, and look twice.**
|
|
1674
|
+
//
|
|
1675
|
+
// This asked once, immediately, and the circumstance it is asked in is
|
|
1676
|
+
// precisely a screen that is still arriving: a content-driven screen whose
|
|
1677
|
+
// rows have not rendered hashes differently from the one memory expected —
|
|
1678
|
+
// which is what produced the `unexpected-screen` — *and* does not yet hold
|
|
1679
|
+
// the next step's target. So the one signal that could have rescued the batch
|
|
1680
|
+
// was read at the only moment it was guaranteed to be absent, and a correct
|
|
1681
|
+
// navigation aborted the run.
|
|
1682
|
+
//
|
|
1683
|
+
// Measured cost of getting this wrong: a field report lost a 7-step plan at
|
|
1684
|
+
// step 5 on a tap that had correctly advanced a wizard, and every discarded
|
|
1685
|
+
// step is ~20s of model latency to re-plan. A settle here costs under a
|
|
1686
|
+
// second and is paid only on the failure path.
|
|
1687
|
+
//
|
|
1688
|
+
// It cannot manufacture a continue: the next step's own target still has to
|
|
1689
|
+
// resolve, which is strong evidence about where we are.
|
|
1690
|
+
for (let look = 0; look < 2; look += 1) {
|
|
1691
|
+
await api.waitFor(deviceQuery, {
|
|
1692
|
+
mode: 'settle', stableMs: 300, timeoutMs: look === 0 ? 1200 : 800, options,
|
|
1693
|
+
}).catch(() => null);
|
|
1694
|
+
try {
|
|
1695
|
+
const hit = await api.locate(deviceQuery, String(target), { refresh: true, options });
|
|
1696
|
+
if (hit?.target) return true;
|
|
1697
|
+
} catch {
|
|
1698
|
+
/* not here yet; one more look, then the halt stands */
|
|
1699
|
+
}
|
|
1567
1700
|
}
|
|
1701
|
+
return false;
|
|
1568
1702
|
}
|
|
1569
1703
|
|
|
1570
1704
|
/**
|
|
@@ -2153,6 +2287,36 @@ export const SWEEP_SECTIONS = 10;
|
|
|
2153
2287
|
|
|
2154
2288
|
const sweepKey = (r) => `${alnum(r.label)}\u0000${Math.round((r.x ?? 0) / 8)}`;
|
|
2155
2289
|
|
|
2290
|
+
/**
|
|
2291
|
+
* Every name a caller may legitimately write for a swept row.
|
|
2292
|
+
*
|
|
2293
|
+
* `sweep` matched on `r.label` alone, and in React Native most interactive
|
|
2294
|
+
* controls have a `testID` and no accessibility label — so a field `sweep`
|
|
2295
|
+
* itself had just printed came back as **"NOT FOUND anywhere"**:
|
|
2296
|
+
*
|
|
2297
|
+
* swept 3 section(s) ... 29 distinct element(s);
|
|
2298
|
+
* NOT FOUND anywhere: "create-service-request-3-requested-by-input"
|
|
2299
|
+
* ...
|
|
2300
|
+
* #18 field 201,480 create-service-request-3-requested-by-input
|
|
2301
|
+
*
|
|
2302
|
+
* Four lines apart, in one response. The next call filled it by `#18` first
|
|
2303
|
+
* try. The reporter called it the most confidence-damaging failure of the run,
|
|
2304
|
+
* because "NOT FOUND anywhere" is a strong claim and it briefly convinced them
|
|
2305
|
+
* the form did not have the field they were looking at.
|
|
2306
|
+
*
|
|
2307
|
+
* **The third time this same assumption has been reported.** Item 152 was the
|
|
2308
|
+
* identifier missing from `matching.rank`; `ee305d4` was a label the matcher
|
|
2309
|
+
* refused; this is `sweep` carrying its own ad-hoc matcher that never learned
|
|
2310
|
+
* either fix. A resolver per call site is a resolver that has to be corrected
|
|
2311
|
+
* per call site — the names belong in one place, which is what this is.
|
|
2312
|
+
*/
|
|
2313
|
+
const sweepNames = (r) => [r.label, r.identifier, ...(r.aliases ?? [])].filter(Boolean).map(alnum);
|
|
2314
|
+
/** Does this row answer to `needle` by any of its names? */
|
|
2315
|
+
const sweepHolds = (r, needle) => {
|
|
2316
|
+
const want = alnum(needle);
|
|
2317
|
+
return want ? sweepNames(r).some((n) => n.includes(want)) : false;
|
|
2318
|
+
};
|
|
2319
|
+
|
|
2156
2320
|
async function sectionHere(deviceQuery, options) {
|
|
2157
2321
|
try {
|
|
2158
2322
|
const map = await view.screenMap(deviceQuery, { options, refresh: true });
|
|
@@ -2298,12 +2462,29 @@ async function sweep(deviceQuery, udid, step, ctx) {
|
|
|
2298
2462
|
// scroll back to it.
|
|
2299
2463
|
if (fill) {
|
|
2300
2464
|
for (const [label, text] of Object.entries(fill)) {
|
|
2301
|
-
if (!here.some((r) =>
|
|
2465
|
+
if (!here.some((r) => sweepHolds(r, label))) continue;
|
|
2302
2466
|
try {
|
|
2303
|
-
|
|
2304
|
-
|
|
2467
|
+
// **Keep what the step said.** `paste` and `type` both read the field
|
|
2468
|
+
// back and say so — "unconfirmed", "reads empty", the quiet-field
|
|
2469
|
+
// caveat — and this threw all of it away and printed a bare
|
|
2470
|
+
// `filled "X"`. A field report called that out as the worst failure
|
|
2471
|
+
// mode an automation tool has: a silent no-op reported as a confirmed
|
|
2472
|
+
// action, caught only because the tester took a screenshot on a
|
|
2473
|
+
// hunch. The honesty existed one function down and stopped here.
|
|
2474
|
+
const body = String(text);
|
|
2475
|
+
// Explicit wins; otherwise length decides. See KEYBOARD_UP_TO.
|
|
2476
|
+
const viaPaste = step.paste === true
|
|
2477
|
+
|| (step.paste !== false && body.length > KEYBOARD_UP_TO);
|
|
2478
|
+
const said = await runStep(deviceQuery, udid, {
|
|
2479
|
+
action: viaPaste ? 'paste' : 'type', into: label, text: body,
|
|
2305
2480
|
}, ctx);
|
|
2306
|
-
|
|
2481
|
+
// Anything the step qualified travels with the claim. A step that
|
|
2482
|
+
// confirmed the read-back says nothing extra, so a clean fill still
|
|
2483
|
+
// reads cleanly.
|
|
2484
|
+
const caveat = /unconfirmed|NOT CONFIRMED|reads empty|did not land|nothing was read back/i.test(said ?? '')
|
|
2485
|
+
? ` — ${String(said).replace(/^(pasted|typed)[^[]*/i, '').trim() || said}`
|
|
2486
|
+
: '';
|
|
2487
|
+
filled.push(`${JSON.stringify(label)} in section ${section + 1}${caveat}`);
|
|
2307
2488
|
} catch (err) {
|
|
2308
2489
|
filled.push(`${JSON.stringify(label)} FAILED in section ${section + 1}: ${err.message.split('\n')[0].slice(0, 90)}`);
|
|
2309
2490
|
}
|
|
@@ -2311,7 +2492,7 @@ async function sweep(deviceQuery, udid, step, ctx) {
|
|
|
2311
2492
|
}
|
|
2312
2493
|
}
|
|
2313
2494
|
|
|
2314
|
-
if (wanted && [...seen.values()].some((r) =>
|
|
2495
|
+
if (wanted && [...seen.values()].some((r) => sweepHolds(r, wanted))) break;
|
|
2315
2496
|
if (fill && !Object.keys(fill).length) break;
|
|
2316
2497
|
prev = here;
|
|
2317
2498
|
await scrollOne(deviceQuery, udid, 'down', ctx);
|
|
@@ -2319,7 +2500,7 @@ async function sweep(deviceQuery, udid, step, ctx) {
|
|
|
2319
2500
|
|
|
2320
2501
|
const all = [...seen.values()];
|
|
2321
2502
|
ctx.sweep = all;
|
|
2322
|
-
const hits = wanted ? all.filter((r) =>
|
|
2503
|
+
const hits = wanted ? all.filter((r) => sweepHolds(r, wanted)) : [];
|
|
2323
2504
|
const listed = (wanted ? hits : all).slice(0, 30)
|
|
2324
2505
|
.map((r) => `[${r.section}] ${JSON.stringify(String(r.label).slice(0, 36))} @${r.x},${r.y}`);
|
|
2325
2506
|
const unfilled = fill ? Object.keys(fill) : [];
|
|
@@ -2334,6 +2515,31 @@ async function sweep(deviceQuery, udid, step, ctx) {
|
|
|
2334
2515
|
+ (listed.length ? `: ${listed.join(', ')}` : '');
|
|
2335
2516
|
}
|
|
2336
2517
|
|
|
2518
|
+
/**
|
|
2519
|
+
* A `waitFor` is satisfied, but the screen is still arriving — hold or go?
|
|
2520
|
+
*
|
|
2521
|
+
* The largest single loss in the 0.15.1 field report: `waitFor "Records"` was
|
|
2522
|
+
* satisfied by a count header reading `78 Records`, the next `tap` fired before
|
|
2523
|
+
* any row had rendered, and the batch aborted with **13 steps unattempted**.
|
|
2524
|
+
* The reporter had even warned us in their own `supervise` text that lists
|
|
2525
|
+
* there render a count header before rows — and the supervisor never got asked,
|
|
2526
|
+
* because from `waitFor`'s point of view the wait had succeeded.
|
|
2527
|
+
*
|
|
2528
|
+
* `stillFillingIn` already recognises exactly this ("a header promises 78
|
|
2529
|
+
* records and only 2 rows are here yet"), and it was wired only to a note on a
|
|
2530
|
+
* step result — advice for the model, costing a round trip to act on. Holding
|
|
2531
|
+
* locally costs milliseconds.
|
|
2532
|
+
*
|
|
2533
|
+
* **It can only ever delay, never fail.** When the budget runs out the wait
|
|
2534
|
+
* still succeeds, carrying what it saw. A `waitFor` that turned a present
|
|
2535
|
+
* target into an error would be the false negative this release just fixed
|
|
2536
|
+
* elsewhere, and a worse trade than the race it is guarding.
|
|
2537
|
+
*/
|
|
2538
|
+
function holdForContent(found, limit) {
|
|
2539
|
+
if (Date.now() >= limit) return null;
|
|
2540
|
+
return stillFillingIn(found?.entry) ?? null;
|
|
2541
|
+
}
|
|
2542
|
+
|
|
2337
2543
|
async function runStep(deviceQuery, udid, step, ctx) {
|
|
2338
2544
|
switch (step.action) {
|
|
2339
2545
|
case 'tap': {
|
|
@@ -2778,8 +2984,12 @@ async function runStep(deviceQuery, udid, step, ctx) {
|
|
|
2778
2984
|
for (const one of alternatives) {
|
|
2779
2985
|
try {
|
|
2780
2986
|
const found = await api.locate(deviceQuery, one, { refresh: step.refresh !== false, options: ctx.options });
|
|
2987
|
+
const arriving = holdForContent(found, limit);
|
|
2988
|
+
if (arriving) { lastError = arriving; continue; }
|
|
2989
|
+
const late = stillFillingIn(found?.entry);
|
|
2781
2990
|
return `${JSON.stringify(one)} appeared at ${found.target.x},${found.target.y}`
|
|
2782
|
-
+ ` (first of ${alternatives.length} awaited)
|
|
2991
|
+
+ ` (first of ${alternatives.length} awaited)`
|
|
2992
|
+
+ (late ? ` [but ${late} — waited out the timeout]` : '');
|
|
2783
2993
|
} catch (err) {
|
|
2784
2994
|
lastError = err.message;
|
|
2785
2995
|
}
|
|
@@ -2819,7 +3029,17 @@ async function runStep(deviceQuery, udid, step, ctx) {
|
|
|
2819
3029
|
// is the price, and it is the same trade: a read is cheaper than the
|
|
2820
3030
|
// round trip a wrong verdict causes.
|
|
2821
3031
|
const found = await api.locate(deviceQuery, query, { index: step.index, refresh: step.refresh !== false, options: ctx.options });
|
|
2822
|
-
|
|
3032
|
+
const arriving = holdForContent(found, limit);
|
|
3033
|
+
if (arriving) {
|
|
3034
|
+
// Still coming. Sleep a beat and look again — the loop re-resolves
|
|
3035
|
+
// fresh, so this cannot lock onto a stale reading.
|
|
3036
|
+
await api.waitFor(deviceQuery, { mode: 'stable', stableMs: 200, timeoutMs: 700, options: ctx.options })
|
|
3037
|
+
.catch(() => null);
|
|
3038
|
+
continue;
|
|
3039
|
+
}
|
|
3040
|
+
const late = stillFillingIn(found?.entry);
|
|
3041
|
+
return `"${found.target.label}" appeared at ${found.target.x},${found.target.y}`
|
|
3042
|
+
+ (late ? ` [but ${late} — waited out the timeout; the screen may still be arriving]` : '');
|
|
2823
3043
|
} catch (err) {
|
|
2824
3044
|
lastError = err.message;
|
|
2825
3045
|
// Waiting cannot make a thing unique.
|
|
@@ -3114,13 +3334,23 @@ export function settleEvidence(w) {
|
|
|
3114
3334
|
*/
|
|
3115
3335
|
export function flowSummary(res, { withTime = true } = {}) {
|
|
3116
3336
|
const time = withTime && Number.isFinite(res.totalMs) ? ` in ${res.totalMs}ms` : '';
|
|
3117
|
-
|
|
3337
|
+
// A step that was allowed not to be there and was not there did not run, and
|
|
3338
|
+
// saying so is the difference between "the nag screen was gone" and "the nag
|
|
3339
|
+
// screen was dismissed". Same lesson as the denominator above: a count that
|
|
3340
|
+
// quietly means two things is the defect.
|
|
3341
|
+
const skipped = (res.results ?? []).filter((r) => r.skipped).length;
|
|
3342
|
+
// After the time, in parentheses. Written inline before it, this read
|
|
3343
|
+
// "skipped as absent in 3412ms" — which says the skipping took 3.4 seconds.
|
|
3344
|
+
const skips = skipped
|
|
3345
|
+
? ` (${skipped} optional step${skipped === 1 ? '' : 's'} skipped as absent)`
|
|
3346
|
+
: '';
|
|
3347
|
+
if (res.ok) return `flow completed — ${res.ranSteps}/${res.totalSteps} steps${time}${skips}`;
|
|
3118
3348
|
const failed = (res.results ?? []).filter((r) => r.ok === false).length || 1;
|
|
3119
3349
|
const worked = Math.max(0, res.ranSteps - failed);
|
|
3120
3350
|
const unattempted = Math.max(0, res.totalSteps - res.ranSteps);
|
|
3121
3351
|
return `FLOW FAILED — ${worked} ok, ${failed} failed`
|
|
3122
3352
|
+ (unattempted ? `, ${unattempted} not attempted` : '')
|
|
3123
|
-
+ ` (of ${res.totalSteps})${time}`;
|
|
3353
|
+
+ ` (of ${res.totalSteps})${time}${skips}`;
|
|
3124
3354
|
}
|
|
3125
3355
|
|
|
3126
3356
|
/**
|
package/src/cli.js
CHANGED
|
@@ -285,6 +285,20 @@ async function main() {
|
|
|
285
285
|
// is a different machine and should say so rather than be guessed at.
|
|
286
286
|
if (flags.readyTimeoutMs) options.readyTimeoutMs = num(flags.readyTimeoutMs);
|
|
287
287
|
|
|
288
|
+
// `--help` on any command, not only as the command.
|
|
289
|
+
//
|
|
290
|
+
// `simframe hpi --help` did the two worst things in sequence: without
|
|
291
|
+
// `--device` it demanded a device *before* printing help, and with one it
|
|
292
|
+
// ignored `--help` and ran the report. A flag that is ignored is worse than
|
|
293
|
+
// one that is rejected — it silently does something other than what was
|
|
294
|
+
// asked. (The general case, unrecognised flags being dropped everywhere
|
|
295
|
+
// rather than here, is filed separately; this is the one that was measured
|
|
296
|
+
// costing a round trip.)
|
|
297
|
+
if (flags.help || flags.h) {
|
|
298
|
+
process.stdout.write(USAGE);
|
|
299
|
+
return;
|
|
300
|
+
}
|
|
301
|
+
|
|
288
302
|
switch (command) {
|
|
289
303
|
case undefined:
|
|
290
304
|
case '-h':
|
|
@@ -977,6 +991,9 @@ async function main() {
|
|
|
977
991
|
return;
|
|
978
992
|
}
|
|
979
993
|
emit(flags, res, [
|
|
994
|
+
...(res.startedElsewhere
|
|
995
|
+
? [`NOTE: recorded starting on ${res.startedElsewhere.recorded}, replayed from ${res.startedElsewhere.here}`]
|
|
996
|
+
: []),
|
|
980
997
|
...(res.results ?? []).map(stepLine),
|
|
981
998
|
actions.flowSummary(res, { withTime: false }),
|
|
982
999
|
]);
|
|
@@ -1242,6 +1259,14 @@ async function main() {
|
|
|
1242
1259
|
? `HPI_time and HPI need a human baseline — none of ${report.overall.flows_measured} measured flow(s) has one yet.`
|
|
1243
1260
|
: `HPI_time ${report.overall.hpi_time} (harmonic mean over ${report.overall.flows_with_human_baseline} flow(s)), HPI ${report.overall.hpi}`,
|
|
1244
1261
|
`step_ratio ${report.overall.step_ratio ?? '—'} (target ≤1.5), model turns per flow ${report.overall.model_turns_median ?? '—'}`,
|
|
1262
|
+
// Printed with what it costs, because the number alone means nothing
|
|
1263
|
+
// to a reader and the whole point is that it is the dominant term in
|
|
1264
|
+
// wall clock — far larger than anything inside simframe.
|
|
1265
|
+
`steps per model call ${report.overall.steps_per_call ?? '—'}`
|
|
1266
|
+
+ (report.overall.steps_per_call
|
|
1267
|
+
? ` — about ${(((20000 + 1700 * report.overall.steps_per_call) / report.overall.steps_per_call) / 1000).toFixed(1)}s`
|
|
1268
|
+
+ ' per step end to end, of which simframe is ~1.7s. Raise this, not the engine.'
|
|
1269
|
+
: ''),
|
|
1245
1270
|
flags.out ? `wrote ${flags.out}` : null,
|
|
1246
1271
|
]);
|
|
1247
1272
|
return;
|
package/src/graph.js
CHANGED
|
@@ -869,7 +869,7 @@ export const VERDICTS = ['ok', 'no-visible-change', 'unexpected-screen', 'unveri
|
|
|
869
869
|
* that already exists do its job. A string still works and still means "exact
|
|
870
870
|
* match only", which is right for a stored prediction that has no tokens.
|
|
871
871
|
*/
|
|
872
|
-
function sameScreen(udid, a, b) {
|
|
872
|
+
export function sameScreen(udid, a, b) {
|
|
873
873
|
const hashOf = (v) => (typeof v === 'string' ? v : v?.hash);
|
|
874
874
|
const ha = hashOf(a);
|
|
875
875
|
const hb = hashOf(b);
|
package/src/matching.js
CHANGED
|
@@ -220,7 +220,49 @@ export function rank(targets, intent, { screen } = {}) {
|
|
|
220
220
|
// them disappears into the ceiling.
|
|
221
221
|
scored.push({ target: t, score, reasons });
|
|
222
222
|
}
|
|
223
|
-
|
|
223
|
+
scored.sort((a, b) => b.score - a.score);
|
|
224
|
+
|
|
225
|
+
// **A distinctive fragment of one long name, when nothing else came close.**
|
|
226
|
+
//
|
|
227
|
+
// Reported from the field: a `waitFor` on a seven-digit record number gave up
|
|
228
|
+
// after 20 s on a screen whose own "Visible:" list printed that number inside
|
|
229
|
+
// a heading — `Record #<digits>`, seventeen characters. The reporter's
|
|
230
|
+
// guess was that `#` was being treated as significant, or that the matcher
|
|
231
|
+
// was anchored. It is neither — the number scores **0.287** against the 0.45
|
|
232
|
+
// floor, because the substring branch in `nameScore` scales by how much of
|
|
233
|
+
// the name the query covers, and seven digits are 41% of that label.
|
|
234
|
+
//
|
|
235
|
+
// That scaling is right and stays: it is what stops "back" beating a real
|
|
236
|
+
// back button from inside a long list row. But its purpose is to resolve
|
|
237
|
+
// *competition*, and when there is no competition it is charging a penalty
|
|
238
|
+
// for a risk that does not exist. So the promotion fires only when **nothing
|
|
239
|
+
// reached the floor** and **exactly one element** contains the query. The
|
|
240
|
+
// "back" case is untouched, because a screen with a back button has at least
|
|
241
|
+
// two names containing "back" and this never runs.
|
|
242
|
+
//
|
|
243
|
+
// The score lands just above the floor, not at 1: it is an act of
|
|
244
|
+
// desperation, not a confident match, and the reason says so — so an
|
|
245
|
+
// ambiguity check downstream still has something honest to weigh.
|
|
246
|
+
if (!scored.some((c) => c.score >= MINIMUM_SCORE)) {
|
|
247
|
+
const q = norm(bare) || norm(intent);
|
|
248
|
+
const holds = (t) => [t.label, t.identifier, ...(t.aliases ?? [])]
|
|
249
|
+
.filter(Boolean)
|
|
250
|
+
.some((n) => norm(n).includes(q) && norm(n) !== q);
|
|
251
|
+
const only = q.length >= 3 ? visible.filter(holds) : [];
|
|
252
|
+
if (only.length === 1) {
|
|
253
|
+
const t = only[0];
|
|
254
|
+
const existing = scored.find((c) => c.target === t);
|
|
255
|
+
const reason = `the only element on this screen containing "${intent}"`;
|
|
256
|
+
if (existing) {
|
|
257
|
+
existing.score = MINIMUM_SCORE + 0.01;
|
|
258
|
+
existing.reasons.push(reason);
|
|
259
|
+
} else {
|
|
260
|
+
scored.push({ target: t, score: MINIMUM_SCORE + 0.01, reasons: [reason] });
|
|
261
|
+
}
|
|
262
|
+
scored.sort((a, b) => b.score - a.score);
|
|
263
|
+
}
|
|
264
|
+
}
|
|
265
|
+
return scored;
|
|
224
266
|
}
|
|
225
267
|
|
|
226
268
|
/** How close two candidates may be before the answer counts as ambiguous. */
|
package/src/mcp.js
CHANGED
|
@@ -13,6 +13,7 @@ import { REGION_COLS, REGION_ROWS, regionMap } from './analyze.js';
|
|
|
13
13
|
import * as actions from './actions.js';
|
|
14
14
|
import * as api from './index.js';
|
|
15
15
|
import * as input from './input.js';
|
|
16
|
+
import * as supervisor from './supervisor.js';
|
|
16
17
|
import * as metrics from './metrics.js';
|
|
17
18
|
import * as navigate from './navigate.js';
|
|
18
19
|
import { bootedDevices, listDevices, permissionServices, resolveDevice } from './platform/index.js';
|
|
@@ -133,7 +134,7 @@ const TOOLS = [
|
|
|
133
134
|
steps: {
|
|
134
135
|
type: 'array',
|
|
135
136
|
description:
|
|
136
|
-
'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
|
|
137
|
+
'Ordered steps. Every selector below accepts "Save" | "#3" | "@120,400", in that order of preference. Act: {"tap":"Save"} (add "index" if a label is ambiguous), {"type":{"into":"Name","text":"Fryer 3"}}, {"paste":{"into":"Notes","text":"long text"}}, {"clear":"Notes"} to empty a field and "clear":true on a type/paste to replace rather than append (drop "into" to type into whatever already has focus, which is how you follow a browser next-field chevron — nothing can be read back then, and the step says so), {"scroll":"down"}, {"scrollTo":"Delete account"}, {"swipe":{"from":[x,y],"to":[x,y]}}, {"button":"HOME"}, {"key":"return"} (the keyboard return/enter key, which is how a mobile search field submits — also escape, tab, space, backspace, and the arrows), {"launch":{"value":"com.example.app","relaunch":true,"args":["-uiTest","1"]}}, {"openUrl":"myapp://x"}, {"permission":{"value":"photos","grant":"grant","bundleId":"com.example.app"}}. Check: {"assert":{"value":"Saved","is":"visible"}} (also gone | enabled | disabled | value with "equals"), {"waitFor":{"value":"Saved","timeoutMs":5000}} (add "failIfStillFor":15000 to stop early once the screen has plainly stopped changing — a 180s wait once burned three minutes on an app that had logged itself out; without it a timeout still reports how long the screen had been still), {"settle":{"stableMs":600}}, {"pause":300}. Recover without a round trip: add "or" to any step for fallback selectors tried locally — {"tap":"Save","or":["Done","Confirm"]} — and "optional":true to any step that may legitimately have nothing to act on, which is how you cross a first-launch nag, a permission sheet or a "What\'s New" in one batch — {"tap":"Not Now","optional":true} is skipped when nothing matches and runs normally when it does, so the rest of the plan survives an interstitial that did not appear — and {"seek":"change username","budget":6} explores for something not on this screen: it OPENS containers (a real action — state changes), checks, and returns to where it started, refusing to open anything that commits, abandons or answers. It does not tap the target; it leaves you on the screen where the target resolves so you tap it next. Do not point it into a flow whose progress you cannot afford to lose. A long screen is only knowable a viewport at a time, so {"sweep":"all","fill":{"Last Name":"Asadi","Email":"a@b.c"}} goes to the top, then reads and fills section by section to the bottom — filling each field while it is on screen, which beats finding one and scrolling back. Add "from":"here" to sweep down from where you are. It reports which section each element was in, what it filled, and what it never found at any scroll position. Prefer it to scrollTo on forms and long lists. Brief the supervisor from the plan: top-level "supervise" is standing guidance for the whole batch ("lists here render a count header before rows; REVIEW stays disabled until a provider is chosen") and per-step "expect" adds to it. When it stops a run the result names the steps it did not attempt — re-issue them with a corrected "supervise" note if the judgement was wrong.',
|
|
137
138
|
items: { type: 'object' },
|
|
138
139
|
},
|
|
139
140
|
autoSettle: {
|
|
@@ -1046,12 +1047,55 @@ async function doScript(target, args, options) {
|
|
|
1046
1047
|
lines.push(`supervisor at step ${s_.index}: ${s_.decision} — ${s_.outcome}`
|
|
1047
1048
|
+ (s_.reason ? ` (it said: "${s_.reason}")` : ''));
|
|
1048
1049
|
}
|
|
1050
|
+
// **A brief with nobody to read it says so.**
|
|
1051
|
+
//
|
|
1052
|
+
// `supervise` is standing guidance for a supervisor; *which* supervisor comes
|
|
1053
|
+
// from `options.supervisor` or SIMFRAME_SUPERVISOR, and with neither set the
|
|
1054
|
+
// brief is discarded. The code even documented that as "always safe" — safe
|
|
1055
|
+
// and silent, which is the trap. Two independent field reports hit it: both
|
|
1056
|
+
// passed a brief on every call, neither ever saw a verdict, and one had
|
|
1057
|
+
// written guidance describing the exact race that then aborted their batch.
|
|
1058
|
+
// Their words: "as written, `supervise` is a prompt I can't observe the
|
|
1059
|
+
// effect of."
|
|
1060
|
+
//
|
|
1061
|
+
// Said only when a brief was actually passed, so it is never noise, and it
|
|
1062
|
+
// names the remedy rather than the condition.
|
|
1063
|
+
if (args.supervise && !supervisor.requested(options)) {
|
|
1064
|
+
lines.push('supervisor: none enabled, so the "supervise" brief was not consulted'
|
|
1065
|
+
+ ' — set supervisor (per call) or SIMFRAME_SUPERVISOR to enable one.');
|
|
1066
|
+
} else if (args.supervise && !(res.supervisions ?? []).length) {
|
|
1067
|
+
lines.push(`supervisor: ${supervisor.requested(options)} enabled, never consulted`
|
|
1068
|
+
+ ' — no step failed in a way that asks for a ruling.');
|
|
1069
|
+
}
|
|
1049
1070
|
if (args.saveAs) {
|
|
1050
1071
|
const saved = navigate.saveFlow(res.device.udid, args.saveAs, res);
|
|
1072
|
+
// Say what to do about it. The refusal this replaced named a condition and
|
|
1073
|
+
// no remedy, and the reporter read it as "recording a flow is impossible"
|
|
1074
|
+
// — which it was.
|
|
1075
|
+
const how = saved.provisional
|
|
1076
|
+
? ' — PROVISIONAL, because this was the first traversal and nothing had been'
|
|
1077
|
+
+ ' seen before to compare against. Replay it once with sim_flow_run and it'
|
|
1078
|
+
+ ' is confirmed. A replay costs no model calls.'
|
|
1079
|
+
: ' — confirmed; replay with sim_flow_run for zero model calls';
|
|
1080
|
+
// One sentence per reason, because "incomplete-run" and "failed-steps" ask
|
|
1081
|
+
// for opposite things from the caller and the bare reason names neither.
|
|
1082
|
+
const why = {
|
|
1083
|
+
'contradicted-steps': 'a step landed somewhere memory says it should not have,'
|
|
1084
|
+
+ ' so this route is not safe to replay. Check the flagged step before re-running it.',
|
|
1085
|
+
'failed-steps': 'a step ran and failed, so replaying this would fail the same way'
|
|
1086
|
+
+ ' every time. Fix the step and re-run the batch, then save.',
|
|
1087
|
+
'incomplete-run': 'the run stopped before its last step, so the rest of the route'
|
|
1088
|
+
+ ' was never observed. Re-run it to the end, then save.',
|
|
1089
|
+
}[saved.reason];
|
|
1090
|
+
// `verdicts` has a hole wherever a step never produced one — a failed step
|
|
1091
|
+
// has no verification — and joining those printed `(unverified, unverified, )`.
|
|
1092
|
+
const seen = (saved.verdicts ?? []).filter(Boolean);
|
|
1051
1093
|
lines.push(
|
|
1052
1094
|
saved.ok
|
|
1053
|
-
? `saved as flow "${saved.name}" (${saved.steps} steps)
|
|
1054
|
-
: `NOT saved as "${args.saveAs}": ${saved.reason}
|
|
1095
|
+
? `saved as flow "${saved.name}" (${saved.steps} steps)${how}`
|
|
1096
|
+
: `NOT saved as "${args.saveAs}": ${saved.reason}`
|
|
1097
|
+
+ `${why ? ` — ${why}` : ''}`
|
|
1098
|
+
+ `${seen.length ? ` (verdicts: ${seen.join(', ')})` : ''}`,
|
|
1055
1099
|
);
|
|
1056
1100
|
}
|
|
1057
1101
|
const escalated = (res.results ?? []).some((r) => metrics.ESCALATING_VERDICTS.has(r.verification?.verdict));
|
|
@@ -1137,6 +1181,14 @@ async function flowRun(target, args, options) {
|
|
|
1137
1181
|
};
|
|
1138
1182
|
}
|
|
1139
1183
|
const lines = [`flow "${args.name}"`, ...stepLines(res)];
|
|
1184
|
+
// Said out loud or not said at all. A flow now records the screen it was
|
|
1185
|
+
// recorded on, and a mismatch that only exists in the return value is a fact
|
|
1186
|
+
// nobody reads — which is the same shape as the check that could not fail.
|
|
1187
|
+
if (res.startedElsewhere) {
|
|
1188
|
+
lines.push(`NOTE: this flow was recorded starting on screen ${res.startedElsewhere.recorded},`
|
|
1189
|
+
+ ` and this replay started on ${res.startedElsewhere.here}. Not refused — content-driven`
|
|
1190
|
+
+ ' screens legitimately change identity — but if the steps below fail to resolve, this is why.');
|
|
1191
|
+
}
|
|
1140
1192
|
lines.push('', await mapFrom(target, options, res.endScreen, { verdictLine: verdictLineFor(res.results) }));
|
|
1141
1193
|
return { content: [text(lines.join('\n'))], isError: !res.ok };
|
|
1142
1194
|
}
|
package/src/metrics.js
CHANGED
|
@@ -654,7 +654,12 @@ export function hpi({ flows, baselines = {} }) {
|
|
|
654
654
|
const agent = quartiles(finished.map((r) => r.wall_time_ms));
|
|
655
655
|
const human = baselines[name]?.wall_time_ms ?? null;
|
|
656
656
|
const humanMedian = human?.p50 ?? null;
|
|
657
|
-
|
|
657
|
+
// Completed runs only, for the same reason the time is: a run that stops
|
|
658
|
+
// three steps in reports a LOW step ratio, so breakage flattered this
|
|
659
|
+
// number too. The runner's suite, with 3 of 14 runs completing, reported
|
|
660
|
+
// `step_ratio 0.375` — "the agent uses a third of the human's steps" about
|
|
661
|
+
// flows that mostly never arrived.
|
|
662
|
+
const stepRatios = finished.map((r) => r.step_ratio).filter((x) => Number.isFinite(x));
|
|
658
663
|
return {
|
|
659
664
|
flow: name,
|
|
660
665
|
runs: runs.length,
|
|
@@ -667,6 +672,22 @@ export function hpi({ flows, baselines = {} }) {
|
|
|
667
672
|
wrong_action: runs.filter((r) => r.wrong_action_taken).length,
|
|
668
673
|
escalations: runs.reduce((acc, r) => acc + (r.escalation_count ?? 0), 0),
|
|
669
674
|
model_turns: median(runs.map((r) => r.model_turns)),
|
|
675
|
+
// **The latency lever, and it was already in the log unsurfaced.**
|
|
676
|
+
//
|
|
677
|
+
// Wall-clock per step is dominated by the model round trip, not by
|
|
678
|
+
// simframe: measured in the field, ~20 s of agent latency against ~1.7 s
|
|
679
|
+
// of simframe work per step. So per-step time is
|
|
680
|
+
// `(20s + 1.7s x n) / n` for n steps in one call — 21.7 s at n=1, 11.7 s
|
|
681
|
+
// at n=2, 6.7 s at n=4. Nothing about making simframe faster moves that;
|
|
682
|
+
// only n does.
|
|
683
|
+
//
|
|
684
|
+
// Which makes this the number to watch, and every hard-fail that drops a
|
|
685
|
+
// caller back to single-stepping a latency regression. Over 186 recorded
|
|
686
|
+
// runs on the author's device the median was 2.0 and the 25th percentile
|
|
687
|
+
// 1.0 — the p25 tail is recovery, and it is where the time goes.
|
|
688
|
+
steps_per_call: median(runs
|
|
689
|
+
.filter((r) => Number.isFinite(r.steps_taken) && r.model_turns > 0)
|
|
690
|
+
.map((r) => r.steps_taken / r.model_turns)),
|
|
670
691
|
};
|
|
671
692
|
}).sort((a, b) => a.flow.localeCompare(b.flow));
|
|
672
693
|
|
|
@@ -684,6 +705,9 @@ export function hpi({ flows, baselines = {} }) {
|
|
|
684
705
|
hpi_accuracy: accuracy,
|
|
685
706
|
hpi_time: hpiTime == null ? null : Number(hpiTime.toFixed(3)),
|
|
686
707
|
hpi: hpiTime == null || accuracy == null ? null : Number((accuracy * hpiTime).toFixed(3)),
|
|
708
|
+
steps_per_call: median(flows
|
|
709
|
+
.filter((f) => Number.isFinite(f.steps_taken) && f.model_turns > 0)
|
|
710
|
+
.map((f) => f.steps_taken / f.model_turns)),
|
|
687
711
|
step_ratio: median(perFlow.map((f) => f.step_ratio)),
|
|
688
712
|
model_turns_median: median(perFlow.map((f) => f.model_turns)),
|
|
689
713
|
},
|
|
@@ -799,8 +823,34 @@ export function breakdown(records, { session = null, flow = null } = {}) {
|
|
|
799
823
|
* than one, and accuracy stays strict at any drop at all. Numbers in
|
|
800
824
|
* docs/BENCHMARKS.md under "What the gate is set to, and why".
|
|
801
825
|
*/
|
|
826
|
+
/**
|
|
827
|
+
* The band HPI_time is *reported* against. It no longer fails a build.
|
|
828
|
+
*
|
|
829
|
+
* It was a gate, at 25% here and written down as 10% in CLAUDE.md — and the
|
|
830
|
+
* measurement cannot support either. `HPI_time` is `human_p50 / agent_p50`
|
|
831
|
+
* where the human was recorded once, on a laptop, and the agent is measured
|
|
832
|
+
* wherever CI happens to run: a hosted runner put the same flows at 37 s and
|
|
833
|
+
* 64 s against 11.5 s and 11.7 s on that laptop, so the ratio mixes the code's
|
|
834
|
+
* speed with the host's. Even on one machine, identical code spans 0.406-0.558
|
|
835
|
+
* across device conditions — a 37% spread. A band inside that can only be
|
|
836
|
+
* silent or wrong, and it was silent: items 148, 152, 154 and 169 all shipped
|
|
837
|
+
* without it firing.
|
|
838
|
+
*
|
|
839
|
+
* So time is a trend, per host, and accuracy is the gate. Accuracy asks
|
|
840
|
+
* whether the agent reached the destination and whether it tapped the wrong
|
|
841
|
+
* thing — properties of the code, not of the machine.
|
|
842
|
+
*/
|
|
802
843
|
export const TIME_REGRESSION = 0.25;
|
|
803
844
|
|
|
845
|
+
/**
|
|
846
|
+
* The most steps-per-human-step we will accept. From the research: an agent
|
|
847
|
+
* taking half again as many actions as a person is wandering.
|
|
848
|
+
*
|
|
849
|
+
* An absolute target rather than a regression, deliberately — it is a ratio of
|
|
850
|
+
* two step counts, so unlike the time it does not change with the host.
|
|
851
|
+
*/
|
|
852
|
+
export const STEP_RATIO_CEILING = 1.5;
|
|
853
|
+
|
|
804
854
|
/**
|
|
805
855
|
* The HPI_time a gate should compare: the median across passes when a report
|
|
806
856
|
* has them, and the single measurement when it does not — so a baseline
|
|
@@ -816,7 +866,7 @@ export const gateTime = (o) => o?.hpi_time_median_of_passes ?? o?.hpi_time ?? nu
|
|
|
816
866
|
* fingerprint eval both shipped unable to fail, and both looked exactly like
|
|
817
867
|
* this. Returns the reasons it should fail — empty means pass.
|
|
818
868
|
*/
|
|
819
|
-
export function gateAgainst(baseline, measured, {
|
|
869
|
+
export function gateAgainst(baseline, measured, { stepRatioCeiling = STEP_RATIO_CEILING } = {}) {
|
|
820
870
|
const base = baseline?.overall ?? {};
|
|
821
871
|
const now = measured?.overall ?? measured ?? {};
|
|
822
872
|
const baseTime = gateTime(base);
|
|
@@ -825,19 +875,44 @@ export function gateAgainst(baseline, measured, { timeRegression = TIME_REGRESSI
|
|
|
825
875
|
if (base.hpi_accuracy != null && now.hpi_accuracy != null && now.hpi_accuracy < base.hpi_accuracy) {
|
|
826
876
|
failures.push(`HPI_accuracy dropped: ${now.hpi_accuracy} < ${base.hpi_accuracy} (any drop fails)`);
|
|
827
877
|
}
|
|
828
|
-
|
|
878
|
+
// Steps, which the host cannot inflate: a ratio of two step counts.
|
|
879
|
+
if (now.step_ratio != null && now.step_ratio > stepRatioCeiling) {
|
|
829
880
|
failures.push(
|
|
830
|
-
`
|
|
881
|
+
`step_ratio ${now.step_ratio} is above the ${stepRatioCeiling} ceiling — the agent is taking more actions than a person`,
|
|
831
882
|
);
|
|
832
883
|
}
|
|
833
|
-
//
|
|
834
|
-
//
|
|
884
|
+
// **HPI_time does not fail a build.** See TIME_REGRESSION. It is reported
|
|
885
|
+
// with its band so a trend is visible and a big move is obvious to a human,
|
|
886
|
+
// and `timeTrend` below is what prints it.
|
|
887
|
+
//
|
|
888
|
+
// A checkout missing the human baseline is still a failure, because it is a
|
|
889
|
+
// broken measurement rather than a slow one — a report with no time at all
|
|
890
|
+
// would otherwise pass by having nothing to say.
|
|
835
891
|
if (baseTime != null && nowTime == null) {
|
|
836
892
|
failures.push('HPI_time is null but the baseline has one — the human baseline it needs is missing from this checkout');
|
|
837
893
|
}
|
|
838
894
|
return failures;
|
|
839
895
|
}
|
|
840
896
|
|
|
897
|
+
/**
|
|
898
|
+
* HPI_time as a line to read, never a verdict.
|
|
899
|
+
*
|
|
900
|
+
* Says how far it moved and whether that is outside the reporting band, and
|
|
901
|
+
* says plainly that it is not a gate — a CI line that looks like a threshold
|
|
902
|
+
* gets read as one.
|
|
903
|
+
*/
|
|
904
|
+
export function timeTrend(baseline, measured, { timeRegression = TIME_REGRESSION } = {}) {
|
|
905
|
+
const baseTime = gateTime(baseline?.overall ?? {});
|
|
906
|
+
const nowTime = gateTime(measured?.overall ?? measured ?? {});
|
|
907
|
+
if (baseTime == null || nowTime == null) return 'HPI_time — not comparable (one side has no measurement)';
|
|
908
|
+
const delta = (nowTime - baseTime) / baseTime;
|
|
909
|
+
const pct = `${delta >= 0 ? '+' : ''}${(delta * 100).toFixed(1)}%`;
|
|
910
|
+
const wide = Math.abs(delta) > timeRegression;
|
|
911
|
+
return `HPI_time ${baseTime} -> ${nowTime} (${pct})`
|
|
912
|
+
+ `${wide ? ` — outside the ${timeRegression * 100}% reporting band, worth a look` : ''}`
|
|
913
|
+
+ ' — reported, not gated: this ratio mixes the code with the host';
|
|
914
|
+
}
|
|
915
|
+
|
|
841
916
|
/** A flow id that sorts by time and is short enough to read in a log. */
|
|
842
917
|
export function newFlowId() {
|
|
843
918
|
return `${Date.now().toString(36)}-${Math.random().toString(36).slice(2, 8)}`;
|
package/src/navigate.js
CHANGED
|
@@ -139,18 +139,69 @@ export function knownScreens(udid) {
|
|
|
139
139
|
return graph.allNodes(udid).map((n) => ({ name: graph.describe(n) ?? n.hash.slice(0, 8), hash: n.hash.slice(0, 8), edges: n.edges.length }));
|
|
140
140
|
}
|
|
141
141
|
|
|
142
|
+
/**
|
|
143
|
+
* A verdict that is evidence *against* a step, as opposed to no evidence yet.
|
|
144
|
+
*
|
|
145
|
+
* This distinction is the whole of the fix below. `unexpected-*` means the step
|
|
146
|
+
* did something other than what memory predicted — a real objection. `unverified`
|
|
147
|
+
* means this edge has never been walked before, which on a first traversal is
|
|
148
|
+
* true of every step by definition and says nothing about whether it worked.
|
|
149
|
+
*/
|
|
150
|
+
const CONTRADICTED = /^unexpected/;
|
|
151
|
+
|
|
142
152
|
/**
|
|
143
153
|
* Save a flow.
|
|
144
154
|
*
|
|
145
|
-
*
|
|
146
|
-
*
|
|
147
|
-
*
|
|
155
|
+
* The rule was "only flows that verified end to end", and it was right about
|
|
156
|
+
* replay safety and wrong about arithmetic: **a first successful traversal is
|
|
157
|
+
* all-`unverified` by construction**, so nothing could ever be recorded, so
|
|
158
|
+
* `sim_flow_run` was unreachable. A field report hit it on a clean 10-of-10
|
|
159
|
+
* batch — *"I never obtained a saved flow, so `sim_flow_run` went untested"* —
|
|
160
|
+
* and it matters far more than its severity suggests: a replayed flow costs
|
|
161
|
+
* **zero model calls**, which is the only path to human-level wall clock. A
|
|
162
|
+
* gate nobody can pass protects nothing and blocks the fastest thing here.
|
|
163
|
+
*
|
|
164
|
+
* So the refusal now needs evidence against a step, not the absence of evidence
|
|
165
|
+
* for it. A flow that ran every step with nothing contradicted saves as
|
|
166
|
+
* **provisional**, and one clean replay promotes it — a bootstrap in two runs,
|
|
167
|
+
* where the second run is the confirmation and is useful anyway.
|
|
168
|
+
*
|
|
169
|
+
* Still refused, because these are real objections: any `unexpected-*` verdict,
|
|
170
|
+
* and a flow that did not reach its own last step.
|
|
148
171
|
*/
|
|
149
172
|
export function saveFlow(udid, name, script, { force = false } = {}) {
|
|
150
|
-
const
|
|
151
|
-
|
|
152
|
-
|
|
173
|
+
const results = script.results ?? [];
|
|
174
|
+
const verdicts = results.map((r) => r.verification?.verdict);
|
|
175
|
+
const contradicted = verdicts.filter((v) => v && CONTRADICTED.test(v));
|
|
176
|
+
const ranAll = script.ranSteps == null
|
|
177
|
+
|| script.steps == null
|
|
178
|
+
|| script.ranSteps >= (script.steps?.length ?? 0);
|
|
179
|
+
// A step that ran and failed is not a step that ran.
|
|
180
|
+
//
|
|
181
|
+
// `ranSteps` counts steps *attempted*, and a failing step stops the batch —
|
|
182
|
+
// so a run whose failure is on the **last** step has `ranSteps === steps.length`
|
|
183
|
+
// and passed the gate above. A peer watched `FLOW FAILED — 4 ok, 1 failed (of
|
|
184
|
+
// 5)` save itself as a flow, which is a flow guaranteed to fail forever. The
|
|
185
|
+
// predicate wanted "every step succeeded" and was written as "every step was
|
|
186
|
+
// reached"; on every run except one they are the same sentence.
|
|
187
|
+
//
|
|
188
|
+
// `!r.ok` rather than `r.ok === false` on purpose: a result that does not say
|
|
189
|
+
// it succeeded has not said it succeeded, and a gate that only catches an
|
|
190
|
+
// explicit `false` is one absent field away from being unable to fail.
|
|
191
|
+
const failedSteps = results.filter((r) => !r.ok);
|
|
192
|
+
if (!force && contradicted.length) {
|
|
193
|
+
return { ok: false, reason: 'contradicted-steps', verdicts };
|
|
194
|
+
}
|
|
195
|
+
if (!force && failedSteps.length) {
|
|
196
|
+
return { ok: false, reason: 'failed-steps', verdicts, failed: failedSteps.map((r) => r.index ?? null) };
|
|
153
197
|
}
|
|
198
|
+
if (!force && !ranAll) {
|
|
199
|
+
return { ok: false, reason: 'incomplete-run', verdicts };
|
|
200
|
+
}
|
|
201
|
+
// Provisional whenever any step lacked a confirming verdict. Recorded on the
|
|
202
|
+
// flow rather than inferred later, so a replay can promote it and a listing
|
|
203
|
+
// can say which flows are still on their first observation.
|
|
204
|
+
const provisional = verdicts.some((v) => !v || v !== 'ok');
|
|
154
205
|
const dir = flowDir(udid);
|
|
155
206
|
fs.mkdirSync(dir, { recursive: true });
|
|
156
207
|
const body = {
|
|
@@ -158,9 +209,28 @@ export function saveFlow(udid, name, script, { force = false } = {}) {
|
|
|
158
209
|
savedAt: Date.now(),
|
|
159
210
|
steps: script.steps ?? (script.results ?? []).map((r) => r.step).filter(Boolean),
|
|
160
211
|
startScreen: script.startScreen ?? null,
|
|
212
|
+
...(provisional ? { provisional: verdicts.filter(Boolean) } : {}),
|
|
161
213
|
};
|
|
162
214
|
store.writeAtomic(path.join(dir, `${encodeURIComponent(name)}.json`), JSON.stringify(body, null, 2));
|
|
163
|
-
return { ok: true, name, steps: body.steps.length };
|
|
215
|
+
return { ok: true, name, steps: body.steps.length, provisional };
|
|
216
|
+
}
|
|
217
|
+
|
|
218
|
+
/**
|
|
219
|
+
* A provisional flow that replays cleanly becomes a confirmed one.
|
|
220
|
+
*
|
|
221
|
+
* The second half of the bootstrap. Without it "provisional" would be a label
|
|
222
|
+
* that never comes off, and the honest state of a flow that has now worked
|
|
223
|
+
* twice is "confirmed".
|
|
224
|
+
*/
|
|
225
|
+
export function confirmFlow(udid, name) {
|
|
226
|
+
const flow = loadFlow(udid, name);
|
|
227
|
+
if (!flow?.provisional) return false;
|
|
228
|
+
const { provisional, ...rest } = flow;
|
|
229
|
+
store.writeAtomic(
|
|
230
|
+
path.join(flowDir(udid), `${encodeURIComponent(name)}.json`),
|
|
231
|
+
JSON.stringify({ ...rest, confirmedAt: Date.now() }, null, 2),
|
|
232
|
+
);
|
|
233
|
+
return true;
|
|
164
234
|
}
|
|
165
235
|
|
|
166
236
|
export function loadFlow(udid, name) {
|
|
@@ -174,7 +244,18 @@ export function listFlows(udid) {
|
|
|
174
244
|
.filter((f) => f.endsWith('.json'))
|
|
175
245
|
.map((f) => store.readJson(path.join(dir, f)))
|
|
176
246
|
.filter(Boolean)
|
|
177
|
-
|
|
247
|
+
// `provisional` travels with the listing, for two reasons. A caller should
|
|
248
|
+
// be able to see which of their flows are still on a first observation
|
|
249
|
+
// without opening the file — and without it the integration check that
|
|
250
|
+
// asserts promotion reads `undefined`, which is always falsy, so the check
|
|
251
|
+
// would pass whether or not anything was promoted. A vacuous check is worse
|
|
252
|
+
// than no check.
|
|
253
|
+
.map((f) => ({
|
|
254
|
+
name: f.name,
|
|
255
|
+
steps: f.steps?.length ?? 0,
|
|
256
|
+
savedAt: f.savedAt,
|
|
257
|
+
...(f.provisional ? { provisional: true } : {}),
|
|
258
|
+
}));
|
|
178
259
|
}
|
|
179
260
|
|
|
180
261
|
export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}) {
|
|
@@ -185,6 +266,28 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
|
|
|
185
266
|
// a human doing the same thing. `minSteps` comes from the flow definition or
|
|
186
267
|
// stays null — the step count of a recorded route is not a claim about the
|
|
187
268
|
// shortest one.
|
|
269
|
+
// Is this the screen the flow was recorded on?
|
|
270
|
+
//
|
|
271
|
+
// Reported, not refused, and the distinction is deliberate. A replay from the
|
|
272
|
+
// wrong screen is mostly self-limiting — the first selector does not resolve
|
|
273
|
+
// and the run stops one step in, and the destructive vocabulary is barred by
|
|
274
|
+
// the verify barrier either way. Refusing on a mismatch would put the exact
|
|
275
|
+
// failure mode of item 174 — content-driven screens fragmenting into several
|
|
276
|
+
// identities — in front of the one path that costs zero model calls. So this
|
|
277
|
+
// says what it saw and lets the run proceed, which also measures how often
|
|
278
|
+
// the mismatch is spurious. If it turns out to be rare, it can become a gate;
|
|
279
|
+
// deciding that by reasoning is how 174 got built in the first place.
|
|
280
|
+
let startedElsewhere = null;
|
|
281
|
+
if (flow.startScreen?.hash) {
|
|
282
|
+
try {
|
|
283
|
+
const here = await api.screenIdentity(device.udid, {});
|
|
284
|
+
if (here.hash && !graph.sameScreen(device.udid, here, flow.startScreen)) {
|
|
285
|
+
startedElsewhere = { recorded: flow.startScreen.hash.slice(0, 8), here: here.hash.slice(0, 8) };
|
|
286
|
+
}
|
|
287
|
+
} catch {
|
|
288
|
+
/* not knowing where we are is not a reason to refuse to try */
|
|
289
|
+
}
|
|
290
|
+
}
|
|
188
291
|
const result = await runScript(device.udid, {
|
|
189
292
|
steps: flow.steps,
|
|
190
293
|
stopOnUnexpected: true,
|
|
@@ -192,5 +295,18 @@ export async function runFlow(deviceQuery, name, { options, ...runOptions } = {}
|
|
|
192
295
|
minSteps: flow.minSteps ?? null,
|
|
193
296
|
...runOptions,
|
|
194
297
|
});
|
|
195
|
-
|
|
298
|
+
// Same arithmetic as `saveFlow`, and the same defect: reaching the last step
|
|
299
|
+
// is not passing it. `result.ok` is the run's own verdict and was ignored
|
|
300
|
+
// here, so a replay that failed on its final assert promoted the flow it had
|
|
301
|
+
// just disproved. Both of a peer's saved flows were marked confirmed by
|
|
302
|
+
// replays that failed; the word "clean" in "confirmed by one clean replay"
|
|
303
|
+
// did not exist in the code path.
|
|
304
|
+
const everyStepPassed = (result.results ?? []).every((r) => r.ok);
|
|
305
|
+
const ok = result.ok !== false && everyStepPassed && result.ranSteps === flow.steps.length;
|
|
306
|
+
// A clean replay is the confirmation a first traversal could not give.
|
|
307
|
+
// Only when nothing was contradicted — a replay that ran to the end while
|
|
308
|
+
// objecting to a step is not a promotion.
|
|
309
|
+
const objected = (result.results ?? []).some((r) => CONTRADICTED.test(r.verification?.verdict ?? ''));
|
|
310
|
+
const promoted = ok && !objected ? confirmFlow(device.udid, name) : false;
|
|
311
|
+
return { ok, name, ...result, ...(promoted ? { promoted: true } : {}), ...(startedElsewhere ? { startedElsewhere } : {}) };
|
|
196
312
|
}
|