@swmansion/argent 0.15.0 → 0.15.1-next.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -2
- package/bin/argent-android-devtools-0.1.0.apk +0 -0
- package/bin/darwin/ax-service +0 -0
- package/bin/darwin/tcp/ax-service +0 -0
- package/bin/darwin/tvos-ax-service +0 -0
- package/bin/darwin/tvos-hid-daemon +0 -0
- package/dist/cli-cmds.mjs +610 -81
- package/dist/cli.js +29 -11
- package/dist/cli.js.map +1 -1
- package/dist/installer-help.d.ts +88 -0
- package/dist/installer-help.js +191 -0
- package/dist/installer-help.js.map +1 -0
- package/dist/installer.mjs +284 -246
- package/dist/mcp-server.mjs +296 -63
- package/dist/preview-ui/index.html +69 -0
- package/dist/tool-server.cjs +50205 -42930
- package/dylibs/libArgentInjectionBootstrap.dylib +0 -0
- package/dylibs/libKeyboardPatch.dylib +0 -0
- package/dylibs/libNativeDevtoolsIos.dylib +0 -0
- package/dylibs/tcp/libArgentInjectionBootstrap.dylib +0 -0
- package/dylibs/tcp/libKeyboardPatch.dylib +0 -0
- package/dylibs/tcp/libNativeDevtoolsIos.dylib +0 -0
- package/dylibs/tvos/libArgentInjectionBootstrap.dylib +0 -0
- package/dylibs/tvos/libKeyboardPatch.dylib +0 -0
- package/dylibs/tvos/libNativeDevtoolsIos.dylib +0 -0
- package/package.json +1 -1
- package/skills/argent-create-flow/SKILL.md +137 -106
- package/skills/argent-device-interact/SKILL.md +3 -3
- package/skills/argent-native-profiler/SKILL.md +1 -0
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
package/package.json
CHANGED
|
@@ -3,143 +3,174 @@ name: argent-create-flow
|
|
|
3
3
|
description: Record a reusable flow (scripted sequence of MCP tool calls) that can be replayed later with a single command. Use when the user asks to create, record, or build a flow, or to script a sequence of device actions.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
##
|
|
6
|
+
## Overview
|
|
7
7
|
|
|
8
|
-
A flow is a
|
|
8
|
+
A flow is a sequence of steps saved to a `.yaml` file in the `.argent/flows/` directory. Each recorded step is **executed live** as you add it, so you verify it works before it becomes part of the flow. Replay a finished flow with `flow-execute`, or — for an e2e flow — headlessly with `argent flow run <name>`.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
Flows store **no device id**: the runner binds a device (the single booted one, or pass `device`/`platform`). A recorded coordinate `gesture-tap` is captured as a portable `tap: { selector }` step whenever the tapped element has stable text/identifier.
|
|
11
11
|
|
|
12
|
-
|
|
13
|
-
| ------------------------ | -------------------------------------------------------------------------- |
|
|
14
|
-
| `flow-start-recording` | Start recording — takes a name and executionPrerequisite, creates the file |
|
|
15
|
-
| `flow-add-step` | Execute a tool call live and record it if it succeeds |
|
|
16
|
-
| `flow-add-echo` | Add a label/comment that prints during replay |
|
|
17
|
-
| `flow-finish-recording` | Stop recording and get a summary |
|
|
18
|
-
| `flow-read-prerequisite` | Read a flow's execution prerequisite without running it |
|
|
19
|
-
| `flow-execute` | Replay a saved flow by name |
|
|
12
|
+
**Two flow types**
|
|
20
13
|
|
|
21
|
-
|
|
14
|
+
- **e2e** — begins with a `launch:` step, which starts that app from scratch (terminate + relaunch), so the flow controls its own start state. No `executionPrerequisite`. May `run:` fragments; cannot itself be a `run:` target. Record one by adding a `restart-app` of the app under test as the **first** step — it is captured as the `launch` step.
|
|
15
|
+
- **fragment** — doesn't begin with a launch; runs against the device's current state. May declare an `executionPrerequisite` (a documented entry-state contract). Invoked from other flows via a `run:` step, or directly by you at any time.
|
|
22
16
|
|
|
23
|
-
|
|
17
|
+
Both run via `argent flow run <name>` — a fragment simply runs against whatever is on screen (its prerequisite is printed as a reminder). Only e2e flows are meaningful CI/suite entries, since only they give a deterministic verdict from a clean start.
|
|
24
18
|
|
|
25
|
-
|
|
26
|
-
2. **Build step-by-step**: For each action, call `flow-add-step` with the tool name and args. The tool runs immediately — check the result before moving on.
|
|
27
|
-
3. **Add labels**: Use `flow-add-echo` between steps to describe what each section does.
|
|
28
|
-
4. **Finish**: Call `flow-finish-recording` to stop recording. It returns the file path where the flow was saved and a summary of all steps. You can edit the `.yaml` file directly afterwards to remove, reorder, or tweak steps.
|
|
19
|
+
### Step directives
|
|
29
20
|
|
|
30
|
-
|
|
21
|
+
Beyond raw `tool:` steps and `echo:`, flows support declarative directives interpreted by the runner (they are **not** agent-callable tools). **Every directive hard-stops the flow on failure**; later steps are reported `skip`.
|
|
31
22
|
|
|
32
|
-
|
|
23
|
+
| Directive | YAML | Meaning |
|
|
24
|
+
| ----------- | -------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- |
|
|
25
|
+
| `launch` | `- launch: com.acme.app` or `- launch: { ios: …, android: … }` | start the app from scratch (terminate + relaunch) and wait until ready |
|
|
26
|
+
| `tap` | `- tap: Login` or `- tap: { x: 0.5, y: 0.57 }` | tap an element by selector (auto-waits), or a raw normalized point |
|
|
27
|
+
| `type` | `- type: { into: email, text: "a@b.com" }` | focus a field, type, then press Enter to submit + dismiss the keyboard |
|
|
28
|
+
| `scroll-to` | `- scroll-to: "Order #1234"` (scrolls down) or `- scroll-to: { target: …, direction: right, within: … }` | momentum-free scroll until the target is visible |
|
|
29
|
+
| `await` | `- await: { visible: Home }` | wait for a UI condition |
|
|
30
|
+
| `wait` | `- wait: 500` | pause for a fixed number of milliseconds (last resort — prefer `await`) |
|
|
31
|
+
| `assert` | `- assert: { visible: Welcome }` | check a condition, hard-fail if it never holds |
|
|
32
|
+
| `snapshot` | `- snapshot: home` or `- snapshot: { name: home, maxMismatch: 0.5 }` | diff a screenshot against a stored baseline |
|
|
33
|
+
| `run` | `- run: login` | execute a fragment's steps inline |
|
|
33
34
|
|
|
34
|
-
|
|
35
|
+
### Selectors
|
|
35
36
|
|
|
36
|
-
|
|
37
|
-
2. You can also call `flow-read-prerequisite` beforehand to inspect the prerequisite without triggering a run.
|
|
38
|
-
3. Once you pass `prerequisiteAcknowledged: true`, the flow runs all steps in order and returns every tool call result (including screenshots) merged into a single response.
|
|
37
|
+
A **selector** is `{ text?, id?, role? }` (all-must-match; `text`/`role` are case-insensitive substrings, `id` matches the element's testID / accessibilityIdentifier / resource-id exactly, case-insensitive, also accepting the unqualified Android resource-id name — `submit` matches `com.example.app:id/submit`) — the same semantics `await-ui-element` uses, though that tool spells the `id` field `identifier` (flow YAML also accepts `identifier` as an alias for `id`, but `id` is the canonical spelling and what the recorder writes). A bare string is a _loose_ selector: it resolves **identifier-first, then falls back to text** (label/value), so `tap: Login` matches a `testID="Login"` or, failing that, visible text "Login" — no need to know which. Loose fallback applies uniformly to every selector slot (`tap`, `type.into`, `await`, `assert`, `scroll-to`). Use the map form to be strict: `{ id: submit-btn }` (identifier only) or `{ text: Login }` (text only, no fallback).
|
|
39
38
|
|
|
40
|
-
|
|
39
|
+
Selectors resolve against the **full native hierarchy** (iOS: the UIView tree; Android: the complete accessibility hierarchy including not-important views) — strictly more than `describe` or the raw `await-ui-element` tool see (both use the trimmed tree), with complete `testID`/`resource-id` coverage. So an `id` selector works even when `describe` collapses or omits the element — don't fall back to coordinate taps just because a testID isn't visible in `describe` output. And when several elements match — a container aggregates its descendants' text, so `text: "Inner"` matches the wrapping containers too — the action directives (`tap`, `type`, `scroll-to`) pick the **most specific** match: an exact text/identifier match beats a substring hit, then the smallest frame wins.
|
|
41
40
|
|
|
42
|
-
|
|
41
|
+
**Quote strings YAML would mangle.** An unquoted `#` starts a YAML comment — `tap: Order #1234` silently parses as `tap: Order` — and bare `yes`/`no`/`on`/`off`/numbers coerce to non-strings. When a selector or typed text contains `#`, `:`, quotes, or could read as a boolean/number, wrap it: `tap: "Order #1234"`.
|
|
43
42
|
|
|
44
|
-
|
|
43
|
+
### `await` and `assert`
|
|
45
44
|
|
|
45
|
+
The **condition is the key**, and its value is the selector:
|
|
46
|
+
|
|
47
|
+
- `{ visible: Home }`, `{ exists: { id: row } }`, `{ hidden: spinner }`
|
|
48
|
+
- `{ text: { in: <selector>, contains: "Taps:" } }` or `{ text: { in: <selector>, equals: "Taps: 0" } }` — `text` locates an element (`in`) and checks its rendered content against exactly one of `contains` (case-insensitive substring) or `equals` (case-insensitive exact match — use it when boundaries matter: `contains: "Taps: 3"` is also satisfied by "Taps: 30"). Reach for `text` only when the locator is an identifier/role; to assert a string is simply on screen, prefer `{ visible: "Taps: 0" }`.
|
|
49
|
+
- A container's text aggregates its descendants' text (space-joined), so `text` can assert what a testID wrapper visibly shows even when the string lives in a child node. That also means `equals` against a wrapper must match _everything_ it shows or exactly the wrapper's own label/value — targeting the leaf holding exactly the value (or using `contains`) stays the clearer spelling.
|
|
50
|
+
|
|
51
|
+
This condition-as-key form is the only spelling. `await` also accepts an optional `timeout` sibling key in milliseconds — `- await: { visible: Home, timeout: 15000 }` — for a transition that legitimately needs longer than the default budget. **Omit `timeout` by default**: the default budget covers normal transitions, and a habitual generous override just delays failure reporting on every broken step. Add one only after a step demonstrably needs it — it timed out at the default and the wait is legitimately slow (a cold start, a network round-trip, a long animation). `assert` has no timeout override: a check that needs seconds to become true is a wait — spell it `await`.
|
|
52
|
+
|
|
53
|
+
For a custom poll interval or bundleId, drop to an explicit `- tool: await-ui-element` step — but the raw tool polls the trimmed `describe` tree, so a testID it reports as not found can still resolve fine as an `await:` directive (see Selectors). Prefer the directive.
|
|
54
|
+
|
|
55
|
+
### `type` and `scroll-to`
|
|
56
|
+
|
|
57
|
+
`type` presses Enter after typing to commit the value and dismiss the keyboard, so it can't cover later targets. For a chained form whose fields feed one explicit submit — e.g. email then password then a `tap: "Log in"` — set `submit: false` on the intermediate fields so a premature Enter doesn't fire the form early: `type: { into: password, text: "hunter2", submit: false }`.
|
|
58
|
+
|
|
59
|
+
`scroll-to` takes an optional `direction` (`up` | `down` | `left` | `right`, default `down` — so the common case is just `- scroll-to: <selector>`) and optionally a `within: <selector>` that anchors the scroll inside a specific container — required to drive a **nested** scroller (e.g. a horizontal carousel inside a vertical list), since the device can't be asked which container to scroll. It scrolls in bounded momentum-free increments, re-checks after each, and stops if a scroll reveals nothing new (end of the container). `tap`/`type` do **not** scroll — add a `scroll-to` before any target that may be off-screen. It's a no-op when the target is already visible, so a defensive `scroll-to` costs nothing on replay and keeps the flow working on smaller screens.
|
|
60
|
+
|
|
61
|
+
### TV targets (Vega)
|
|
62
|
+
|
|
63
|
+
A Vega (Fire TV) device is remote-driven — there is no touch input, so the touch directives (`tap`, `type`, `scroll-to`) fail on it with guidance. Drive focus with `tool: tv-remote` steps and type with `tool: keyboard` instead; everything else (`launch`, `await`, `assert`, `wait`, `snapshot`, `echo`, `run`, selectors) works unchanged — the tree comes from the on-device automation toolkit, which attaches at app launch (the `launch` step waits for it, so a leading `launch` also guarantees selectors resolve).
|
|
64
|
+
|
|
65
|
+
```yaml
|
|
66
|
+
steps:
|
|
67
|
+
- launch: com.example.app.main # the interactive component id from manifest.toml
|
|
68
|
+
- await: { visible: Home }
|
|
69
|
+
- tool: tv-remote
|
|
70
|
+
args: { button: [down, select] } # move focus, then confirm — one step per navigation
|
|
71
|
+
- await: { visible: Explore Screen }
|
|
72
|
+
- snapshot: explore
|
|
46
73
|
```
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
74
|
+
|
|
75
|
+
Since a `tv-remote` path is positional (like a coordinate tap), gate each navigation with an `await` on the destination screen and echo where focus should be — that is what makes the flow diagnosable when the focus order changes.
|
|
76
|
+
|
|
77
|
+
### Standalone runner
|
|
78
|
+
|
|
79
|
+
`argent flow run <name> [--device <id>] [--platform ios|android|chromium|vega] [--update-baselines] [--output <dir>] [--json]` runs a flow with no LLM in the loop and exits non-zero on any failure — suitable for CI (e2e flows; a fragment runs against the current device state, useful while authoring). `snapshot` baselines live in `.argent/flows/__baselines__/<flow>/`, keyed by platform + resolution; a `snapshot` step **fails** when no baseline exists for the run's device class, so seed baselines with `--update-baselines`, review them, and commit `__baselines__/` — and pin the device class in CI (`--device`/`--platform`, same simulator model) so runs compare against the committed key. The status bar is pinned (iOS `simctl status_bar`, Android demo mode) for the run so it doesn't drive visual diffs. `--output <dir>` writes each failed snapshot's baseline/current/diff images to `<dir>/<flow>/` — a stable path for CI artifact upload.
|
|
80
|
+
|
|
81
|
+
## Tools
|
|
82
|
+
|
|
83
|
+
| Tool | Purpose |
|
|
84
|
+
| ------------------------ | --------------------------------------------------------------------------------------------------------- |
|
|
85
|
+
| `flow-start-recording` | Start recording — takes a name and (fragments only) an optional `executionPrerequisite`; creates the file |
|
|
86
|
+
| `flow-add-step` | Execute a tool call live and record it if it succeeds |
|
|
87
|
+
| `flow-add-echo` | Add a label/comment that prints during replay |
|
|
88
|
+
| `flow-finish-recording` | Stop recording and get a summary |
|
|
89
|
+
| `flow-read-prerequisite` | Read a flow's execution prerequisite without running it |
|
|
90
|
+
| `flow-execute` | Replay a saved flow by name |
|
|
91
|
+
|
|
92
|
+
Every tool during recording returns the current flow file contents, so you can track what has been recorded. Rules:
|
|
93
|
+
|
|
94
|
+
- **Every step runs live.** You see the real tool result (including screenshots) — verify the step worked before continuing. **Only successful steps are recorded**: a failed call writes nothing to the flow file; fix the issue and try again.
|
|
95
|
+
- **Pass `project_root` once.** Give the absolute `project_root` (an error is returned if the path is not absolute) to `flow-start-recording` — it is stored for the session and used by all subsequent flow tools. You do **not** pass a flow name to `flow-add-step`, `flow-add-echo`, or `flow-finish-recording` — the active flow is tracked automatically.
|
|
96
|
+
- **Start before adding.** Calling those tools without an active recording returns _"No active flow. Call flow-start-recording first."_
|
|
97
|
+
- **One flow at a time.** `flow-start-recording` while already recording switches to the new flow — the response tells you which flow was abandoned and which is now active; the old flow's file remains on disk.
|
|
98
|
+
- **Mistakes can be edited out.** Edit the `.yaml` file directly to remove or reorder steps.
|
|
99
|
+
|
|
100
|
+
### flow-add-step arguments
|
|
101
|
+
|
|
102
|
+
The `command` parameter is the MCP tool name; `args` is a **JSON string** (not an object), omitted entirely for tools with no arguments:
|
|
50
103
|
|
|
51
104
|
```
|
|
52
105
|
command: "gesture-tap"
|
|
53
106
|
args: "{\"udid\": \"<UDID>\", \"x\": 0.5, \"y\": 0.35}"
|
|
54
|
-
```
|
|
55
107
|
|
|
56
|
-
```
|
|
57
|
-
command: "screenshot"
|
|
58
|
-
args: "{\"udid\": \"<UDID>\"}"
|
|
59
|
-
```
|
|
60
|
-
|
|
61
|
-
```
|
|
62
108
|
command: "await-ui-element"
|
|
63
109
|
args: "{\"udid\": \"<UDID>\", \"condition\": \"visible\", \"selector\": {\"text\": \"Continue\"}}"
|
|
64
110
|
```
|
|
65
111
|
|
|
66
|
-
Record an `await-ui-element` step to **gate** the next step on a screen transition — it blocks until the element is `visible`/`hidden` (or contains `text`), so the following step runs only once the screen has actually settled
|
|
112
|
+
Record an `await-ui-element` step to **gate** the next step on a screen transition — it blocks until the element is `visible`/`hidden` (or contains `text`), so the following step runs only once the screen has actually settled; prefer this over a fixed `delayMs`. If its condition is not met before the timeout, replay **stops at that step** (the steps after it assume the transition happened). See the `await-ui-element` section of `argent-device-interact` for the full condition/selector reference. The live call sees only the trimmed `describe` tree — if it can't find an identifier you know exists, gate on visible text to get the step recorded, then retarget the identifier in the `await:` form during polish (the directive resolves the full hierarchy — see Selectors); don't conclude the testID is unusable in the flow.
|
|
67
113
|
|
|
68
|
-
|
|
114
|
+
## Recording
|
|
69
115
|
|
|
70
|
-
|
|
116
|
+
1. **Start, then launch as the first step (e2e) or set the stage yourself (fragment).** Call `flow-start-recording` with a descriptive name and the absolute `project_root`. For an **e2e** flow, record a `restart-app` of the app under test as the **first** step — it runs live (resetting the device for the rest of the recording) and is captured as the flow's `launch` step. For a **fragment**, bring the device to the entry state _before_ recording and pass an `executionPrerequisite` describing it (e.g. "App on the login screen") to `flow-start-recording` instead.
|
|
117
|
+
2. **Build step-by-step**: for each action, call `flow-add-step` with the tool name and args. The tool runs immediately — check the result before moving on, and gate each navigation with an `await-ui-element` step.
|
|
118
|
+
3. **Add labels**: use `flow-add-echo` between steps — echo the expected state, not just the action (see _Making flows resilient_).
|
|
119
|
+
4. **Finish**: call `flow-finish-recording`. It returns the file path where the flow was saved and a summary of all steps.
|
|
120
|
+
5. **Polish**: **read the saved `.yaml` file** and convert the raw `tool:` steps that have a cleaner directive form (the recorder leaves these as tools):
|
|
121
|
+
- `tool: keyboard` typing into a field → `type: { into: "<field>", text: "…" }`, folding in the `tap` that focused the field.
|
|
122
|
+
- `tool: await-ui-element` gating a transition → `await: { visible: "…" }` / `{ hidden: … }` / `{ text: { in: …, equals: … } }`, carrying a custom `timeoutMs` over as a `timeout` sibling key. Converting also upgrades the wait from the trimmed `describe` tree to the flow's full-hierarchy tree (see Selectors). Keep the raw `tool: await-ui-element` step only when it sets a custom `pollIntervalMs`/`bundleId` the directive can't express.
|
|
123
|
+
- A scroll-to-reach-an-element — a `tool: gesture-swipe` used to bring a specific element on screen before interacting with it (a `tap`, `type`, `assert`, …) → `scroll-to: { target: "<that element>", direction: … }`, dropping the swipe. This is far more robust than a fixed-distance swipe: it scrolls momentum-free and stops exactly when the target appears, so it survives layout and content changes. (`tap`/`type` do not scroll, so a raw swipe whose fling lands differently on another device leaves the following tap unresolved — always prefer the `scroll-to` rewrite.) Keep a `gesture-swipe` as a raw `tool:` step when it isn't scrolling toward a specific element — especially a velocity-dependent gesture like swipe-to-dismiss, edge-swipe-back, or swipe-to-reveal a row action, which a momentum-free `scroll-to` would not reproduce.
|
|
71
124
|
|
|
72
|
-
-
|
|
73
|
-
- **Only successful steps are recorded.** If a tool call fails, nothing is written to the flow file — fix the issue and try again.
|
|
74
|
-
- **Pass `project_root` only to `flow-start-recording`.** It is stored for the session and automatically used by all subsequent flow tools. An error is returned if the path is not absolute.
|
|
75
|
-
- **You do NOT need to pass a flow name** to `flow-add-step`, `flow-add-echo`, or `flow-finish-recording`. The active flow is tracked automatically after `flow-start-recording`.
|
|
76
|
-
- **Start before adding.** Calling `flow-add-step`, `flow-add-echo`, or `flow-finish-recording` without an active recording returns an error: _"No active flow. Call flow-start-recording first."_
|
|
77
|
-
- **One flow at a time.** If you call `flow-start-recording` while already recording, the active flow switches to the new one. The response tells you which flow was abandoned and which is now active. The old flow's file remains on disk.
|
|
78
|
-
- **Mistakes can be edited out.** If a step was recorded by mistake, edit the `.yaml` file directly to remove or reorder entries.
|
|
125
|
+
Every other recorded tool (`gesture-swipe`, `gesture-scroll`, `button`, `screenshot`, …) has no directive form — leave it as a `tool:` step. The recorder already handles the rest: coordinate `gesture-tap`s are captured as portable `tap:` selector steps, a `restart-app` is captured as a `launch:` step, a `flow-execute` of a sibling fragment is captured as a `run: <name>` composition directive, and device ids are stripped. Captured selectors are emitted in the strict map form (`tap: { text: General }`), never as a loose bare string — the recorder verified the exact element the tap hit, and a bare string would re-parse as loose and route through the identifier-first fallback it was never checked against. After editing, re-run with `flow-execute` to confirm the cleaned flow still passes.
|
|
79
126
|
|
|
80
|
-
|
|
127
|
+
### Example session
|
|
81
128
|
|
|
82
129
|
```
|
|
83
|
-
flow-start-recording { name: "open-
|
|
84
|
-
flow-add-echo { message: "
|
|
85
|
-
flow-add-step { command: "
|
|
86
|
-
flow-add-echo { message: "
|
|
87
|
-
flow-add-step { command: "gesture-tap", args: "{\"udid\": \"ABC\", \"x\": 0.5, \"y\": 0.35}" }
|
|
88
|
-
flow-add-
|
|
130
|
+
flow-start-recording { name: "open-about", project_root: "/Users/dev/MyApp" }
|
|
131
|
+
flow-add-echo { message: "Start Settings from scratch" }
|
|
132
|
+
flow-add-step { command: "restart-app", args: "{\"udid\": \"ABC\", \"bundleId\": \"com.apple.Preferences\"}" } # ⇒ captured as `- launch: com.apple.Preferences` — this is now an e2e flow
|
|
133
|
+
flow-add-echo { message: "On the Settings root list, tapping the 'General' row" }
|
|
134
|
+
flow-add-step { command: "gesture-tap", args: "{\"udid\": \"ABC\", \"x\": 0.5, \"y\": 0.35}" } # ⇒ captured as `- tap: { text: General }` (portable selector, no udid)
|
|
135
|
+
flow-add-step { command: "await-ui-element", args: "{\"udid\": \"ABC\", \"condition\": \"visible\", \"selector\": {\"text\": \"About\"}}" } # gate the transition
|
|
136
|
+
flow-add-echo { message: "On Settings > General, tapping 'About'" }
|
|
89
137
|
flow-add-step { command: "gesture-tap", args: "{\"udid\": \"ABC\", \"x\": 0.5, \"y\": 0.17}" }
|
|
138
|
+
flow-add-step { command: "await-ui-element", args: "{\"udid\": \"ABC\", \"condition\": \"visible\", \"selector\": {\"text\": \"Model Name\"}}" }
|
|
90
139
|
flow-finish-recording {}
|
|
91
140
|
```
|
|
92
141
|
|
|
93
|
-
|
|
142
|
+
Then polish the saved file: the two `await-ui-element` steps become `await:` directives (see the file below).
|
|
94
143
|
|
|
95
|
-
|
|
96
|
-
flow-execute { name: "open-settings", project_root: "/Users/dev/MyApp" }
|
|
97
|
-
→ Returns: notice with executionPrerequisite: "Simulator booted with app installed"
|
|
98
|
-
"Verify the prerequisite is met and call flow-execute again with prerequisiteAcknowledged set to true."
|
|
144
|
+
## Replaying
|
|
99
145
|
|
|
100
|
-
flow-execute
|
|
101
|
-
|
|
102
|
-
|
|
146
|
+
Call `flow-execute` with the flow name (and `project_root`, unless a recording this session already stored it). If the flow has an execution prerequisite, the tool returns a **notice** with the prerequisite text instead of running — verify the prerequisite is met (you can also inspect it beforehand with `flow-read-prerequisite`) and call `flow-execute` again with `prerequisiteAcknowledged: true`. A flow without a prerequisite runs immediately. The run executes all steps in order and returns a structured report: `{ ok, passed, failed, skipped, errored, steps }`.
|
|
147
|
+
|
|
148
|
+
**What each step reports.** Raw `tool:` steps include the underlying tool's full `result` (screenshots and other outputs render as usual). The directive steps are summarized: `tap`/`type`/`await`/`assert` report only `status` + `reason`, and `snapshot` adds `artifacts` only when there is something to look at — a failed comparison (baseline/current/diff paths), a missing-baseline failure (`current` only), or a baseline write; a clean pass reports just `status` + `reason`. So converting a `tool: gesture-tap` into a `tap:` directive during cleanup drops only that tap's (uninteresting) raw result — output-bearing tools like `screenshot` have no directive form and stay `tool:` steps, so their results keep flowing through.
|
|
103
149
|
|
|
104
|
-
##
|
|
150
|
+
## Flow file format
|
|
105
151
|
|
|
106
|
-
|
|
152
|
+
The top-level is an object with `steps` (array) and — fragments only — `executionPrerequisite` (an e2e flow, one beginning with `launch:`, has none). Besides the directives above:
|
|
107
153
|
|
|
108
|
-
- `- echo: <message>` — a label
|
|
109
|
-
- `- tool: <name>` with optional `args:` — a tool call. A tool step may also carry `delayMs: <ms>` to sleep that long before it runs. (`await-ui-element` is an ordinary tool step; see
|
|
154
|
+
- `- echo: <message>` — a label printed during replay
|
|
155
|
+
- `- tool: <name>` with optional `args:` — a raw tool call. A tool step may also carry `delayMs: <ms>` to sleep that long before it runs. (`await-ui-element` is an ordinary tool step; see _flow-add-step arguments_ and _Making flows resilient_ for when to gate a transition with one.)
|
|
110
156
|
|
|
111
|
-
|
|
157
|
+
The polished result of the example session above:
|
|
112
158
|
|
|
113
159
|
```yaml
|
|
114
|
-
executionPrerequisite: Simulator booted with app installed
|
|
115
160
|
steps:
|
|
116
|
-
- echo:
|
|
117
|
-
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
- echo:
|
|
122
|
-
-
|
|
123
|
-
|
|
124
|
-
udid: ABC
|
|
125
|
-
condition: visible
|
|
126
|
-
selector:
|
|
127
|
-
text: General
|
|
128
|
-
- echo: Tap General
|
|
129
|
-
- tool: gesture-tap
|
|
130
|
-
args:
|
|
131
|
-
udid: ABC
|
|
132
|
-
x: 0.5
|
|
133
|
-
y: 0.35
|
|
134
|
-
- echo: Tap About
|
|
135
|
-
- tool: gesture-tap
|
|
136
|
-
args:
|
|
137
|
-
udid: ABC
|
|
138
|
-
x: 0.5
|
|
139
|
-
y: 0.17
|
|
161
|
+
- echo: Start Settings from scratch
|
|
162
|
+
- launch: com.apple.Preferences
|
|
163
|
+
- echo: On the Settings root list, tapping the 'General' row
|
|
164
|
+
- tap: { text: General }
|
|
165
|
+
- await: { visible: About }
|
|
166
|
+
- echo: On Settings > General, tapping 'About'
|
|
167
|
+
- tap: { text: About }
|
|
168
|
+
- await: { visible: Model Name }
|
|
140
169
|
```
|
|
141
170
|
|
|
142
|
-
|
|
171
|
+
Note there is **no device id** anywhere in the file — the recorder strips them and the runner injects the bound device.
|
|
172
|
+
|
|
173
|
+
## When to proactively record a flow
|
|
143
174
|
|
|
144
175
|
You do not need the user to ask for a flow. Record one proactively when you recognize any of these patterns:
|
|
145
176
|
|
|
@@ -148,28 +179,28 @@ You do not need the user to ask for a flow. Record one proactively when you reco
|
|
|
148
179
|
- **Complex path discovered**: You worked through a non-trivial sequence of taps/swipes/navigation to reach a desired app state. Capture it before it is lost.
|
|
149
180
|
- **User says "again" / "one more time"**: Any request to redo what you just did is a signal to record first, then replay.
|
|
150
181
|
|
|
151
|
-
##
|
|
182
|
+
## Flow self-improvement
|
|
152
183
|
|
|
153
184
|
Flows break. UI layouts change, coordinates drift, screens get added or removed. When `flow-execute` returns a failure, follow this procedure to diagnose and fix the flow instead of silently re-recording or giving up.
|
|
154
185
|
|
|
155
|
-
###
|
|
186
|
+
### Classify the result
|
|
156
187
|
|
|
157
188
|
After every `flow-execute`, classify the outcome before proceeding:
|
|
158
189
|
|
|
159
190
|
| Outcome | Signal | Action |
|
|
160
191
|
| ---------------------- | --------------------------------------------------------------------- | ------------------ |
|
|
161
192
|
| **Success** | All steps completed, final screenshot shows expected state | Continue with task |
|
|
162
|
-
| **Hard error** | A step has `ERROR` in the result — engine stopped there | Enter
|
|
163
|
-
| **Silent misfire** | All steps completed but final screenshot shows wrong screen | Enter
|
|
164
|
-
| **Partial divergence** | Intermediate screenshot shows wrong state even though later steps ran | Enter
|
|
193
|
+
| **Hard error** | A step has `ERROR` in the result — engine stopped there | Enter **Diagnose** |
|
|
194
|
+
| **Silent misfire** | All steps completed but final screenshot shows wrong screen | Enter **Diagnose** |
|
|
195
|
+
| **Partial divergence** | Intermediate screenshot shows wrong state even though later steps ran | Enter **Diagnose** |
|
|
165
196
|
|
|
166
|
-
For silent misfires and partial divergence, echo annotations (
|
|
197
|
+
For silent misfires and partial divergence, echo annotations (see _Making flows resilient_) are your reference for what each screen _should_ look like.
|
|
167
198
|
|
|
168
|
-
###
|
|
199
|
+
### Diagnose
|
|
169
200
|
|
|
170
201
|
1. Note the failure step index and error message (if hard error).
|
|
171
202
|
2. Call `screenshot` to see where the app actually is now.
|
|
172
|
-
3. Call `describe` or `debugger-component-tree` to get the current element tree.
|
|
203
|
+
3. Call `describe` or `debugger-component-tree` to get the current element tree. Remember `describe` shows less than the flow tree — a testID missing from its output can still resolve as a selector (see Selectors).
|
|
173
204
|
4. Compare current state to what the failed step expected. Classify the root cause:
|
|
174
205
|
|
|
175
206
|
| Root cause | Symptoms |
|
|
@@ -182,7 +213,7 @@ For silent misfires and partial divergence, echo annotations (§10.5) are your r
|
|
|
182
213
|
|
|
183
214
|
5. State the diagnosis in one sentence before attempting any correction.
|
|
184
215
|
|
|
185
|
-
###
|
|
216
|
+
### Correct
|
|
186
217
|
|
|
187
218
|
Choose the lightest strategy that fits:
|
|
188
219
|
|
|
@@ -190,7 +221,7 @@ Choose the lightest strategy that fits:
|
|
|
190
221
|
Read `.argent/flows/<flow-name>.yaml`, update the broken step's `x`/`y`, `bundleId`, `text`, or other args. Re-run `flow-execute` to verify.
|
|
191
222
|
|
|
192
223
|
**Strategy 2 — Manual recovery + continue** (timing/transient issues, one-off replay).
|
|
193
|
-
Manually execute the failed step with corrected coordinates from
|
|
224
|
+
Manually execute the failed step with corrected coordinates from the Diagnose step, then manually execute remaining steps. Does not fix the YAML — use only when re-recording is not worth it.
|
|
194
225
|
|
|
195
226
|
**Strategy 3 — Re-record from failure point** (structural changes, new intermediate screens).
|
|
196
227
|
Navigate the app to the state just before the failure point. Call `flow-start-recording` with the same flow name (overwrites). Re-add the working prefix steps via `flow-add-step`, then continue recording new steps from the divergence point. Call `flow-finish-recording`.
|
|
@@ -206,22 +237,22 @@ Reset the app to prerequisite state (`restart-app` + `launch-app`). Record from
|
|
|
206
237
|
- 3+ steps broken, or unclear root cause → Strategy 4
|
|
207
238
|
- Flow used for profiling comparison (must be identical) → Strategy 4
|
|
208
239
|
|
|
209
|
-
###
|
|
240
|
+
### Verify and bound retries
|
|
210
241
|
|
|
211
242
|
After applying a correction, re-run `flow-execute` to verify.
|
|
212
243
|
|
|
213
244
|
- If it succeeds → done. Report what changed (e.g. "Fixed step 4: updated tap coordinates from 0.5,0.35 to 0.5,0.42").
|
|
214
|
-
- If it fails at a **different** step → return to
|
|
245
|
+
- If it fails at a **different** step → return to Diagnose for a second attempt.
|
|
215
246
|
- If this is already the second correction attempt → **stop**. Report the diagnosis to the user and recommend a full re-record or manual investigation.
|
|
216
247
|
|
|
217
248
|
**Hard cap: 2 correction cycles.** Do not enter an unbounded fix loop.
|
|
218
249
|
|
|
219
|
-
###
|
|
250
|
+
### Making flows resilient
|
|
220
251
|
|
|
221
252
|
Apply these when recording new flows to reduce future breakage:
|
|
222
253
|
|
|
223
254
|
- **Echo expected state, not just actions.** Write `"On Settings > General screen, about to tap About"` not `"Tap About"`. During diagnosis these tell you what the screen _should_ look like.
|
|
224
|
-
- **Gate transitions with `await-ui-element`, not fixed delays.** After a tap that triggers a navigation, record an `await-ui-element` step that waits for the next screen's element to be `visible` (or a spinner to be `hidden`) before the following step. This removes the **Timing** failure mode in
|
|
255
|
+
- **Gate transitions with `await-ui-element`, not fixed delays.** After a tap that triggers a navigation, record an `await-ui-element` step that waits for the next screen's element to be `visible` (or a spinner to be `hidden`) before the following step — converted to an `await:` directive during polish. This removes the **Timing** failure mode in Diagnose (the element is in the tree but the tap fired before the screen settled) and is more reliable than `delayMs` or an extra `screenshot`. An unmet wait stops replay at that step, so a mistimed step can never run blind.
|
|
225
256
|
- **Add screenshot steps after critical navigation.** Insert `screenshot` steps after screen transitions. These produce images in the flow result you can inspect during diagnosis.
|
|
226
257
|
- **Write specific executionPrerequisites.** `"App on home tab, user logged in, simulator UDID is <X>"` — not `"App running"`. Verify with `screenshot` + `describe` before acknowledging.
|
|
227
258
|
- **Prefer launch-app / open-url over navigation chains.** Deep links are more resilient to layout changes than tap sequences.
|
|
@@ -169,7 +169,7 @@ Values: `home`, `back`, `power`, `volumeUp`, `volumeDown`, `appSwitch`, `actionB
|
|
|
169
169
|
{ "udid": "<UDID>", "text": "search query", "key": "enter" }
|
|
170
170
|
```
|
|
171
171
|
|
|
172
|
-
Special keys: `enter`, `escape`, `backspace`, `tab`, `space`, `arrow-up`, `arrow-down`, `arrow-left`, `arrow-right`, `f1`–`f12`. Optional: `"delayMs": 100` between keystrokes (default 50ms).
|
|
172
|
+
Special keys: `enter`, `escape`, `backspace`, `tab`, `space`, `arrow-up`, `arrow-down`, `arrow-left`, `arrow-right`, `f1`–`f12`. Optional: `"delayMs": 100` between keystrokes (default 50ms) — applies to the iOS simulator and Chromium; it is ignored on Android phones/tablets (typed via `adb input text`, no per-key cadence), on Vega, and on TV targets.
|
|
173
173
|
|
|
174
174
|
### rotate — Change orientation
|
|
175
175
|
|
|
@@ -188,8 +188,8 @@ Instead of polling `screenshot`/`describe` in a loop, use `await-ui-element` to
|
|
|
188
188
|
```
|
|
189
189
|
|
|
190
190
|
- `condition`: `exists`, `visible`, `hidden`, or `text`.
|
|
191
|
-
- `selector`: `{ text?, identifier?, role? }` — every provided field must match
|
|
192
|
-
- Prefer a **specific** selector. A loose substring can match several elements, and the tool may then key off one you didn't mean: `text` reads the
|
|
191
|
+
- `selector`: `{ text?, identifier?, role? }` — every provided field must match. `text` matches the element's label or value and `role` its element role (e.g. `AXButton`, `button`, `TextView`, `StaticText`), both as case-insensitive substrings; `identifier` matches its accessibility id / resource-id / testID **exactly** (case-insensitive), also accepting the unqualified Android resource-id name (`submit` matches `com.example.app:id/submit`). The synthetic `ROOT` container `describe` prints is never matched, so a `role` like `AXGroup`/`html` won't trivially "match the screen".
|
|
192
|
+
- Prefer a **specific** selector. A loose substring can match several elements, and the tool may then key off one you didn't mean: `text` reads the first **visible** match in **reading order** (top-to-bottom, left-to-right — the same order `describe` lists them, so it's the one you saw first; when no match is visible, the first match overall), while `visible`/`exists` are satisfied by **any** match. Disambiguate with a longer or more exact string, an `identifier`, or a `role` (e.g. pin to a text role like `StaticText` to skip a same-named button). On a `text` timeout the `note` quotes the matched element's text, so you can see which one it landed on.
|
|
193
193
|
- `text` condition also needs `expectedText` (substring the matched element must contain).
|
|
194
194
|
- `hidden` treats a selector that matches **nothing** as already-hidden, so a typo'd selector returns an instant (false) success. Double-check the selector for `hidden` waits — the result `note` flags when the selector never matched any element. (On iOS, if the accessibility backend is down the tree comes back empty; the tool will **not** report `hidden` success off such a degraded read and the `note` surfaces the boot hint instead.)
|
|
195
195
|
- Optional `timeoutMs` (default 5000) and `pollIntervalMs` (default 400).
|
|
@@ -27,6 +27,7 @@ After `native-profiler-analyze` surfaces findings, use `profiler-stack-query` to
|
|
|
27
27
|
- **Hang detected** → `profiler-stack-query` mode=`hang_stacks` for full native call chains → mode=`function_callers` for the suspected function → read native source.
|
|
28
28
|
- **CPU hotspot** → `profiler-stack-query` mode=`thread_breakdown` for per-thread distribution → mode=`function_callers` for the dominant function.
|
|
29
29
|
- **Memory leak** → `profiler-stack-query` mode=`leak_stacks` filtered by `object_type` for responsible frames and libraries.
|
|
30
|
+
- iOS: if leaks come back unattributed (responsible frame `<Call stack limit reached>`), re-run `native-profiler-start` with `malloc_stack_logging: true`. This cold-launches the app with Malloc Stack Logging so leaks carry a real allocation backtrace (responsible frame + library). It restarts the app and adds overhead, so use it only when you need leak attribution — not for CPU/hang passes.
|
|
30
31
|
|
|
31
32
|
After presenting findings, ask the user whether to investigate further, implement fixes, or stop. After applying fixes, always re-profile the same scenario and compare with `profiler-load`. Report honestly whether the target metric improved, regressed, or stayed flat. If the fix showed no net benefit or introduced regressions elsewhere, say so and reconsider.
|
|
32
33
|
|