tamash-playwright 0.10.0-beta.5 → 0.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -3,9 +3,9 @@
3
3
  All notable changes to this project are documented here. Format loosely follows
4
4
  [Keep a Changelog](https://keepachangelog.com/); dates are when each version was published.
5
5
 
6
- ## [0.10.0] - 2026-08-26
6
+ ## [0.10.0] - 2026-08-27
7
7
 
8
- ### Fixed (continued)
8
+ ### Fixed
9
9
 
10
10
  - **`waitFor()` is never sent to the AI.** It's a state check, not an action — a timeout on it can
11
11
  mean a genuinely broken selector, or it can mean the element correctly never reached the expected
@@ -18,6 +18,21 @@ All notable changes to this project are documented here. Format loosely follows
18
18
  up, for a heal that could never have succeeded. Now fails fast with a clear `state-wait-not-healed`
19
19
  stage and zero AI calls; a real action (fill/click/...) on the same kind of broken locator is
20
20
  unaffected.
21
+ - **`claude-subscription` was burning a real, unnecessary amount of extra tokens.** The SDK's
22
+ `effort` option silently defaults to `'high'` ("deep reasoning") when left unset, and this
23
+ provider never set it — a small prompt-complexity increase (the `nearbyRef`/`nearbyText` addition
24
+ below) pushed adaptive thinking to reason unpredictably harder for a task that only ever needs to
25
+ return one line of JSON. Confirmed live via real CI runs: output tokens for the same heal went
26
+ from a 188-445 baseline up to 422-881, inconsistently, run to run. `thinking: { type: 'disabled' }`
27
+ + `effort: 'low'` removes the variability entirely — verified repeatedly at a steady 28-30 output
28
+ tokens, below even the pre-regression baseline, with no change to correctness.
29
+ - **The primary ariaSnapshot capture no longer requests `boxes:true`.** Every node was paying for a
30
+ `[box=x,y,w,h]` annotation that nothing on the text/`ref` path (including the `nearbyRef`/
31
+ `adjacent`-strategy widening logic below) ever reads — it's purely topological. The one real
32
+ consumer (vision's own nearest-candidate lookup) already captures its own separate, fresh
33
+ snapshot, so this is genuinely free: verified live, input tokens dropped ~23% on a large real
34
+ page (3794 → 2937) with no loss of accuracy, and no change on small pages (box overhead scales
35
+ with node count).
21
36
 
22
37
  ### Added
23
38
 
@@ -32,27 +47,17 @@ All notable changes to this project are documented here. Format loosely follows
32
47
  wherever the relevant text actually is, not "excludes the nav"), ~29% on a pair of identical
33
48
  sibling fields that still had to be correctly disambiguated, and a clean, correct fallback when
34
49
  the description doesn't match the page's real text at all.
35
-
36
- ### Fixed
37
-
38
- - **The primary ariaSnapshot capture no longer requests `boxes:true`.** Every node was paying for a
39
- `[box=x,y,w,h]` annotation that nothing on the text/`ref` path (including all of today's
40
- `nearbyRef`/`adjacent`-strategy widening logic) ever reads — it's purely topological. The one
41
- real consumer (vision's own nearest-candidate lookup) already captures its own separate, fresh
42
- snapshot, so this is genuinely free: verified live, input tokens dropped ~23% on a large real
43
- page (3794 → 2937) with no loss of accuracy, and no change on small pages (box overhead scales
44
- with node count).
45
- - **`claude-subscription` was burning a real, unnecessary amount of extra tokens.** The SDK's
46
- `effort` option silently defaults to `'high'` ("deep reasoning") when left unset, and this
47
- provider never set it — a small prompt-complexity increase (below) pushed adaptive thinking to
48
- reason unpredictably harder for a task that only ever needs to return one line of JSON.
49
- Confirmed live via real CI runs: output tokens for the same heal went from a 188-445 baseline up
50
- to 422-881, inconsistently, run to run. `thinking: { type: 'disabled' }` + `effort: 'low'`
51
- removes the variability entirely — verified repeatedly at a steady 28-30 output tokens, below
52
- even the pre-regression baseline, with no change to correctness.
53
-
54
- ### Added
55
-
50
+ - **A new `adjacent` selector strategy**, fixing a real ambiguity in the existing `near` strategy:
51
+ when two fields with no identity of their own share a row/section (two dropdowns side by side,
52
+ say), `near`'s "climb to a shared ancestor, then search it for any element of this role"
53
+ approach matches both and gives up rather than risk the wrong one. The AI's `ref` response can
54
+ now optionally report `nearbyRef`/`nearbyText`/`nearbyRole` for a nameless target it identified;
55
+ `deriveDurableLocator` uses that hint to find the true common ancestor between the target and
56
+ its label — via each ref's own full ancestor chain, not by assuming either sits at a matching
57
+ depth — and, when they're proven to be immediate sibling branches, builds a precise CSS
58
+ `:text() + *` sibling match (or an xpath climb-then-step, when the label text turns out to be
59
+ nested below its own branch root). Verified live against a real configured provider, resolving
60
+ the correct field and never its same-row neighbor in both directions.
56
61
  - **Full attempt-history logging**: `SelfHealingReport` now carries an `attempts[]` array — one
57
62
  entry per cache/ref/text/vision/action-recovery attempt actually made, each with its own
58
63
  `succeeded`/`stage`/`error`. Previously only the *last* attempt's stage survived; an earlier
@@ -66,17 +71,6 @@ All notable changes to this project are documented here. Format loosely follows
66
71
  - **`ariaSnapshot` is attached to the test report on failure** — the exact accessibility tree the
67
72
  AI reasoned over, so a confusing report can be diagnosed against real evidence instead of a
68
73
  separately-captured DevTools screenshot.
69
- - **A new `adjacent` selector strategy**, fixing a real ambiguity in the existing `near` strategy:
70
- when two fields with no identity of their own share a row/section (two dropdowns side by side,
71
- say), `near`'s "climb to a shared ancestor, then search it for any element of this role"
72
- approach matches both and gives up rather than risk the wrong one. The AI's `ref` response can
73
- now optionally report `nearbyRef`/`nearbyText`/`nearbyRole` for a nameless target it identified;
74
- `deriveDurableLocator` uses that hint to find the true common ancestor between the target and
75
- its label — via each ref's own full ancestor chain, not by assuming either sits at a matching
76
- depth — and, when they're proven to be immediate sibling branches, builds a precise CSS
77
- `:text() + *` sibling match (or an xpath climb-then-step, when the label text turns out to be
78
- nested below its own branch root). Verified live against a real configured provider, resolving
79
- the correct field and never its same-row neighbor in both directions.
80
74
 
81
75
  ## [0.9.0] - 2026-08-26
82
76
 
package/README.md CHANGED
@@ -267,7 +267,7 @@ Beyond a single broken `click`/`fill`/`getByRole` on the main page, all of this
267
267
 
268
268
  - **Popups and extra tabs.** A page opened via `context.newPage()`, `window.open`, or a `target="_blank"` link is just as healing-aware as your main `page` — no manual wrapping needed.
269
269
  - **Elements inside `<iframe>`s.** `page.frameLocator('#my-iframe')` and anything chained off it heals the same way, scoped correctly to the iframe's own document.
270
- - **Most of the Playwright API surface**, not just clicks and fills — `check`, `selectOption`, `dragTo`, `dispatchEvent`, read methods like `textContent`/`getAttribute`/`isChecked`, `screenshot`, and more. Methods that can't be safely healed by guessing a replacement element (`dragTo`, `drop`) are still reported honestly on failure, they're just never silently retried with a different element.
270
+ - **Most of the Playwright API surface**, not just clicks and fills — `check`, `selectOption`, `dragTo`, `dispatchEvent`, read methods like `textContent`/`getAttribute`/`isChecked`, `screenshot`, and more. A few are deliberately excluded, but still reported honestly on failure rather than silently retried: `dragTo`/`drop` can't be safely healed by guessing a replacement element for just one side of a two-sided drag, and `waitFor` is a state check rather than an action — a timeout on it can mean a genuinely broken selector, or that the element correctly never reached the expected state (verifying something does *not* appear, say), which can't be told apart from the error alone. (`expect(locator).toBeVisible()` and similar assertions are excluded for the same reason, but never even reach this mechanism — they're Playwright's own matcher, not a method this wraps.)
271
271
 
272
272
  ## When there's no name to match: finding elements by structure
273
273
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tamash-playwright",
3
- "version": "0.10.0-beta.5",
3
+ "version": "0.10.0",
4
4
  "description": "Plug and Play Self-healing for Playwright and automatically recovers broken selectors using an AI model (Ollama, OpenAI, Anthropic, Gemini, or a Claude/GitHub Copilot subscription).",
5
5
  "main": "dist/index.js",
6
6
  "types": "dist/index.d.ts",
package/usage.md CHANGED
@@ -161,7 +161,7 @@ Everything below applies identically in both places — same package, same behav
161
161
 
162
162
  - **Popups and extra tabs.** A page opened via `context.newPage()`, `window.open`, or a `target="_blank"` link is just as healing-aware as your main `page`.
163
163
  - **Elements inside `<iframe>`s.** `page.frameLocator('#my-iframe')` and anything chained off it heals the same way, correctly scoped to the iframe's own document.
164
- - **Most of the Playwright API surface**, not just clicks and fills — `check`, `selectOption`, `dragTo`, `dispatchEvent`, read methods like `textContent`/`getAttribute`/`isChecked`, `screenshot`, and more. `dragTo` and `drop` are reported honestly on failure rather than guessed at.
164
+ - **Most of the Playwright API surface**, not just clicks and fills — `check`, `selectOption`, `dragTo`, `dispatchEvent`, read methods like `textContent`/`getAttribute`/`isChecked`, `screenshot`, and more. `dragTo`, `drop`, and `waitFor` are reported honestly on failure rather than guessed at — the first two because a two-sided drag can't be safely fixed by guessing one side, `waitFor` because a timeout on it can be a correct test outcome (e.g. verifying an absence), not a broken selector. `expect(locator).toBeVisible()` and similar assertions are excluded the same way, but never reach this mechanism at all — they're Playwright's own matcher, not a method this wraps.
165
165
 
166
166
  ## When there's no name to match: finding elements by structure
167
167