@aarwitz/tapp 0.17.22 → 0.17.23

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "tapp",
3
3
  "description": "Give Claude hands and eyes on iOS, Android, and web apps, with exploration, replayable flows, evidence, and deterministic CI gates.",
4
- "version": "0.17.22",
4
+ "version": "0.17.23",
5
5
  "author": {
6
6
  "name": "Aaron Horowitz",
7
7
  "url": "https://github.com/aarwitz"
@@ -24,7 +24,7 @@
24
24
  "command": "npx",
25
25
  "args": [
26
26
  "-y",
27
- "@aarwitz/tapp@0.17.22",
27
+ "@aarwitz/tapp@0.17.23",
28
28
  "mcp"
29
29
  ],
30
30
  "cwd": "${CLAUDE_PROJECT_DIR}"
package/AGENTS.md CHANGED
@@ -68,6 +68,7 @@ installs, returns the bundle id) → `tapp_explore {appBundleId}`.
68
68
  | "Find/reach this named screen or control" | `tapp_session_start` with `focus`, or `tapp_focus`; plain CLI: `tapp focus` | screenshot-by-screenshot wandering |
69
69
  | "Tap through / drive / fill a form / log in" | `tapp_session_start` → `session_act` loop | repeated `open_app` calls (cold relaunch each time) |
70
70
  | "Is my app broken? Find bugs" | `tapp_explore` — `appBundleId` for iOS, `androidAppId` for Android, `url` for owned web apps; returns an observation (findings + evidence), not a ship verdict — gate a merge with the CI gate (`tapp ci` CLI / the GitHub Action) + a contract | a manual session (exploration is autonomous) |
71
+ | "Is this live site / production / a prospect's site broken?" | `tapp_audit` (CLI `tapp audit <url> --pages N`) — read-only: dead controls, broken images/links/assets, failed requests, JS errors, mixed content, overflow; writes a capture with evidence | `tapp_explore` (it clicks, types and submits) |
71
72
  | "Make this flow a repeatable test" | drive it in a session, then `tapp_flow_save`; replay with `tapp_flow_run` | re-driving it by hand every time |
72
73
  | "What's on screen right now?" | `tapp_screenshot` / `tapp_ui_tree` | relaunching the app |
73
74
 
@@ -209,3 +210,5 @@ without a coding agent, model, subscription, or API key. AI generation and `asse
209
210
  of retrying variations.
210
211
  - When you show a screenshot as proof, say what it proves and what it doesn't ("login works;
211
212
  I haven't verified checkout").
213
+ - A clean `tapp_audit` means the page is served without structural defects. It says nothing about
214
+ behaviour, because nothing was clicked; never report it as "the site works".
@@ -2512,6 +2512,19 @@ const server = new Server(
2512
2512
  prompts: {},
2513
2513
  tools: {},
2514
2514
  },
2515
+ // Shown to the model by MCP clients at connect time: the routing rules that keep an agent from
2516
+ // reaching for the wrong tool, in the order the mistakes actually happen.
2517
+ instructions: [
2518
+ "Tapp gives you hands and eyes on real iOS, Android and web app surfaces. Pick the smallest operation:",
2519
+ "- see or screenshot one screen → tapp_open_app; what is on screen now → tapp_screenshot / tapp_ui_tree;",
2520
+ "- reach a named screen or drive a journey → tapp_session_start (+ focus) then session_act;",
2521
+ "- find bugs in an app or environment you own → tapp_explore (it clicks, types and submits; minutes);",
2522
+ "- check production or a site you do not own → tapp_audit only (read-only: never clicks);",
2523
+ "- a merge decision → the CI gate (tapp ci / GitHub Action), never exploration.",
2524
+ "Results are observations with findings, coverage and evidence paths — never a score or ship verdict.",
2525
+ "inconclusive:true is not a pass; say what blocked coverage. A clean audit says the page is served without structural defects, not that the product works.",
2526
+ "Open returned screenshot paths with your image tool before describing a screen. Never claim a flow works that you did not drive or see.",
2527
+ ].join("\n"),
2515
2528
  }
2516
2529
  );
2517
2530
 
@@ -2742,7 +2755,7 @@ server.setRequestHandler(ListToolsRequestSchema, async () => ({
2742
2755
  "verdict/releaseScore. Web separates deterministic findings from advisory sampled control probes. " +
2743
2756
  "There is a coverage floor: if the app barely explored (crash on launch / sign-in wall) it reports " +
2744
2757
  "`inconclusive` — absence of findings is NEVER a pass. For iOS the app must already be installed on a booted simulator (use tapp_list_simulators / " +
2745
- "tapp_boot_simulator first). For web, only point it at an app/environment you own — it CLICKS things. " +
2758
+ "tapp_boot_simulator first). For web, only point it at an app/environment you own — it CLICKS things; for production or a site you do not own use tapp_audit (read-only). " +
2746
2759
  "Tapp explores autonomously and does NOT pause to prompt for input — " +
2747
2760
  "it fills forms with safe defaults. The result includes `inputFieldsEncountered` (and `inputHint`): if " +
2748
2761
  "the app showed login/form fields and the user hasn't given you values, ASK THE USER what to enter (offer " +
@@ -668,7 +668,9 @@ export async function exploreWeb({ url, maxActions = 40, timeoutSec = 300, outDi
668
668
  // phone run legitimately disagree about which nav links exist, and a baseline diff across
669
669
  // them must be able to say so instead of reporting "resolved".
670
670
  emit("CONTEXT", { ...(String(device || "").trim() ? { device: String(device).trim() } : {}), viewport: captureProfile.viewport, deviceScaleFactor: captureProfile.deviceScaleFactor ?? 1 });
671
- const context = await browser.newContext(captureProfile);
671
+ // Playwright records natively (no ffmpeg transcode needed, unlike simctl's .mov output) —
672
+ // written to a Playwright-chosen filename in outDir, finalized only on context.close().
673
+ const context = await browser.newContext({ ...captureProfile, recordVideo: { dir: outDir, size: captureProfile.viewport } });
672
674
  await installWebListenerTracking(context);
673
675
  if (watch) await installWebWatchUi(context);
674
676
  const page = await context.newPage();
@@ -1127,7 +1129,10 @@ export async function exploreWeb({ url, maxActions = 40, timeoutSec = 300, outDi
1127
1129
  outbound.total = outboundLinks.size;
1128
1130
  outbound.mailtos = mailtoLinks.size;
1129
1131
  if (outboundLinks.size) {
1130
- const auditPage = await context.newPage();
1132
+ // A fresh, unrecorded context: outbound link checks navigate away to third-party sites and
1133
+ // must not inherit the exploration context's recordVideo (noise, not exploration evidence).
1134
+ const auditContext = await browser.newContext(captureProfile);
1135
+ const auditPage = await auditContext.newPage();
1131
1136
  auditPage.setDefaultTimeout(8000);
1132
1137
  const targets = [...outboundLinks.entries()].slice(0, OUTBOUND_LINK_LIMIT);
1133
1138
  outbound.checked = targets.length;
@@ -1153,7 +1158,7 @@ export async function exploreWeb({ url, maxActions = 40, timeoutSec = 300, outDi
1153
1158
  const phrase = webUnavailableShellPhrase(text);
1154
1159
  if (phrase) issue("outbound_unavailable", "medium", `Outbound link returns 200 but shows "${phrase}": ${href.slice(0, 100)}`, meta.screen, href, meta.sourceUrl);
1155
1160
  }
1156
- await auditPage.close().catch(() => {});
1161
+ await auditContext.close().catch(() => {});
1157
1162
  }
1158
1163
  if (mailtoLinks.size) {
1159
1164
  if (process.env.TAPP_ENFORCE_PUBLIC_EGRESS === "1") {
@@ -1181,6 +1186,15 @@ export async function exploreWeb({ url, maxActions = 40, timeoutSec = 300, outDi
1181
1186
  : "frontier-drained";
1182
1187
  emit("COMPLETE", { actions, screens: screenCount, credentialsProvided: !!(testEmail || testPassword), credentialsUsed: loginTried, timedOut, stop, outbound });
1183
1188
  fs.closeSync(markersFd);
1189
+ // Video finalizes on context.close(), before browser.close() tears down the recorder —
1190
+ // same "exploration.<ext>" name summarizeCapture() and the website replay already look for.
1191
+ const video = page.video();
1192
+ await context.close().catch(() => {});
1193
+ if (video) {
1194
+ await video.path()
1195
+ .then((p) => fs.renameSync(p, path.join(outDir, "exploration.webm")))
1196
+ .catch(() => {});
1197
+ }
1184
1198
  await browser.close().catch(() => {});
1185
1199
  }
1186
1200
  return { markersPath, outDir, actions, screens: screenCount, seedRoutes: normalizedSeeds, seedTargets: normalizedTargets };
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@aarwitz/tapp",
3
- "version": "0.17.22",
3
+ "version": "0.17.23",
4
4
  "mcpName": "io.github.aarwitz/tapp",
5
5
  "description": "Let coding agents verify UI changes on real iOS, Android, and web surfaces, then enforce reviewed proof in deterministic CI.",
6
6
  "license": "MIT",
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: tapp
3
- description: Use Tapp to see, drive, explore, and verify real application surfaces on iOS simulators, Android emulators/devices, or the web. Use when a user asks an agent to test an app or UI change, find bugs, inspect or screenshot a screen, exercise a journey, create a replayable flow, gather release evidence, or run the deterministic Tapp gate. Also use when the user mentions Tapp, @aarwitz/tapp, tapp_* tools, .tapp artifacts, or asks whether agent-authored UI actually works.
3
+ description: Use Tapp to see, drive, explore, and verify real application surfaces on iOS simulators, Android emulators/devices, or the web. Use when a user asks an agent to test an app or UI change, find bugs, inspect or screenshot a screen, exercise a journey, create a replayable flow, gather release evidence, run the deterministic Tapp gate, or check a live website, production, or a site they do not own without touching it (read-only audit). Also use when the user mentions Tapp, @aarwitz/tapp, tapp_* tools, .tapp artifacts, or asks whether agent-authored UI actually works.
4
4
  ---
5
5
 
6
6
  # Tapp
@@ -16,7 +16,8 @@ screen or journey works from source inspection alone.
16
16
  | Inspect controls on the current screen | `tree` / `tapp_ui_tree` |
17
17
  | Reach a named screen/control | `focus` / `tapp_focus` (source + observed UI Map fast path) |
18
18
  | Drive a specific journey | MCP session start → focus or act → end |
19
- | Find bugs autonomously | `explore` / `tapp_explore` |
19
+ | Find bugs autonomously (owned app/environment; it clicks) | `explore` / `tapp_explore` |
20
+ | Check production or a site you do not own (never clicks) | `audit` / `tapp_audit` |
20
21
  | Preserve a journey | record and save a Flow; replay it deterministically |
21
22
  | Decide whether a merge passes policy | `ci`; exploration never decides this |
22
23
 
@@ -42,8 +43,7 @@ For a focused request in an already-grounded repository, use the requested targe
42
43
  than starting another broad exploration. If `.tapp/ui-map.json` does not exist yet, ground it once
43
44
  with `init . --explore`; source alone can locate a surface but cannot authorize unobserved taps.
44
45
  Targets may be a repository path, Xcode container, `.app`, iOS bundle id, APK plus Android app id,
45
- or owned HTTP(S) URL. Never explore a third-party web property without authorization: exploration
46
- clicks and types. For production or a site you do not own use `audit` / `tapp_audit` — read-only.
46
+ or owned HTTP(S) URL. Never explore a third-party web property: exploration clicks and types.
47
47
 
48
48
  ## Navigate like a source-connected expert
49
49
 
@@ -11,7 +11,8 @@ npx -y @aarwitz/tapp@latest explore [target]
11
11
  npx -y @aarwitz/tapp@latest focus "Save storefront settings visible above keyboard" [target]
12
12
  npx -y @aarwitz/tapp@latest open [target]
13
13
  npx -y @aarwitz/tapp@latest tree [target] --json
14
- npx -y @aarwitz/tapp@latest audit https://example.com --json # read-only: never clicks, safe against production
14
+ npx -y @aarwitz/tapp@latest audit https://example.com --pages 5 --json # read-only: never clicks; safe on production or third-party sites
15
+ # several URLs or --urls FILE batch sites · --device "iPhone 13" audits the mobile rendering · --json FILE writes the result · exit 1 = defects found
15
16
  npx -y @aarwitz/tapp@latest shot
16
17
  npx -y @aarwitz/tapp@latest report latest
17
18
  npx -y @aarwitz/tapp@latest doctor