sel2pw-scan 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Shan Konduru
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,88 @@
1
+ Metadata-Version: 2.4
2
+ Name: sel2pw-scan
3
+ Version: 0.1.0
4
+ Summary: Offline inventory of Selenium Java test projects for a Playwright migration assessment
5
+ Author: Shan Konduru
6
+ License-Expression: MIT
7
+ Project-URL: Repository, https://github.com/ShanKonduru/sel2pw-scan
8
+ Project-URL: Issues, https://github.com/ShanKonduru/sel2pw-scan/issues
9
+ Project-URL: Changelog, https://github.com/ShanKonduru/sel2pw-scan/blob/main/CHANGELOG.md
10
+ Keywords: selenium,playwright,migration,assessment,inventory,test-automation,java
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Operating System :: OS Independent
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Programming Language :: Python :: 3 :: Only
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Programming Language :: Python :: 3.14
22
+ Classifier: Topic :: Software Development :: Quality Assurance
23
+ Classifier: Topic :: Software Development :: Testing
24
+ Requires-Python: >=3.10
25
+ Description-Content-Type: text/markdown
26
+ License-File: LICENSE
27
+ Provides-Extra: dev
28
+ Requires-Dist: pytest>=8; extra == "dev"
29
+ Dynamic: license-file
30
+
31
+ # sel2pw-scan
32
+
33
+ Offline inventory of a Selenium Java test project, for assessing a migration to Python Playwright.
34
+
35
+ `sel2pw-scan` reads your `.java` files and build files and writes a report: which test framework you use,
36
+ how many tests, page objects and helper classes you have, and how often your code uses Selenium features
37
+ that need deliberate work in a migration (fixed sleeps, frames, alerts, JavaScript calls, data providers and
38
+ so on).
39
+
40
+ It is a scanner only. It does not convert or change any code, uses no AI model and makes no network calls.
41
+
42
+ ## Install
43
+
44
+ Requires Python 3.10 or newer.
45
+
46
+ ```bash
47
+ pip install sel2pw-scan
48
+ ```
49
+
50
+ ## Use
51
+
52
+ ```bash
53
+ sel2pw-scan path/to/selenium-project --out inventory.md
54
+ ```
55
+
56
+ | Option | Effect |
57
+ |---|---|
58
+ | `--out FILE` | Write the report to a file instead of printing it |
59
+ | `--format md` | Default. Summary report, described below |
60
+ | `--format json` | Full detail for tooling, including locator values and annotation attributes |
61
+ | `--version` | Print the version |
62
+
63
+ `python -m sel2pw_scan ...` works the same way.
64
+
65
+ ## What the report contains
66
+
67
+ The Markdown report (`--format md`) lists:
68
+
69
+ - the scanned folder path, detected frameworks (TestNG, JUnit 4, JUnit 5), Selenium version and build files
70
+ - counts of Java files, test classes, page objects, base and utility classes, and test methods (including disabled ones)
71
+ - how often each of 27 Selenium constructs appears, such as `fixed_sleep`, `frames`, `alerts`,
72
+ `javascript_execution`, `data_provider` and `selenium_grid`
73
+ - every test method as `ClassName#method`, with its file, line, assertion count and constructs
74
+ - every class with its role, parent class and the project classes it depends on
75
+
76
+ It does **not** contain test code, locator values, string literals or test data. Read it before sharing, and
77
+ rename anything you consider sensitive, such as class names.
78
+
79
+ The JSON report (`--format json`) adds locator values and annotation attributes, so treat it as internal.
80
+
81
+ ## Limits
82
+
83
+ The scan uses pattern matching, not a full Java parser, so counts can be slightly off for unusual code.
84
+ It finds tests by annotations such as `@Test`; Cucumber `.feature` files and Kotlin or Groovy sources are not read.
85
+
86
+ ## License
87
+
88
+ [MIT](LICENSE)
@@ -0,0 +1,58 @@
1
+ # sel2pw-scan
2
+
3
+ Offline inventory of a Selenium Java test project, for assessing a migration to Python Playwright.
4
+
5
+ `sel2pw-scan` reads your `.java` files and build files and writes a report: which test framework you use,
6
+ how many tests, page objects and helper classes you have, and how often your code uses Selenium features
7
+ that need deliberate work in a migration (fixed sleeps, frames, alerts, JavaScript calls, data providers and
8
+ so on).
9
+
10
+ It is a scanner only. It does not convert or change any code, uses no AI model and makes no network calls.
11
+
12
+ ## Install
13
+
14
+ Requires Python 3.10 or newer.
15
+
16
+ ```bash
17
+ pip install sel2pw-scan
18
+ ```
19
+
20
+ ## Use
21
+
22
+ ```bash
23
+ sel2pw-scan path/to/selenium-project --out inventory.md
24
+ ```
25
+
26
+ | Option | Effect |
27
+ |---|---|
28
+ | `--out FILE` | Write the report to a file instead of printing it |
29
+ | `--format md` | Default. Summary report, described below |
30
+ | `--format json` | Full detail for tooling, including locator values and annotation attributes |
31
+ | `--version` | Print the version |
32
+
33
+ `python -m sel2pw_scan ...` works the same way.
34
+
35
+ ## What the report contains
36
+
37
+ The Markdown report (`--format md`) lists:
38
+
39
+ - the scanned folder path, detected frameworks (TestNG, JUnit 4, JUnit 5), Selenium version and build files
40
+ - counts of Java files, test classes, page objects, base and utility classes, and test methods (including disabled ones)
41
+ - how often each of 27 Selenium constructs appears, such as `fixed_sleep`, `frames`, `alerts`,
42
+ `javascript_execution`, `data_provider` and `selenium_grid`
43
+ - every test method as `ClassName#method`, with its file, line, assertion count and constructs
44
+ - every class with its role, parent class and the project classes it depends on
45
+
46
+ It does **not** contain test code, locator values, string literals or test data. Read it before sharing, and
47
+ rename anything you consider sensitive, such as class names.
48
+
49
+ The JSON report (`--format json`) adds locator values and annotation attributes, so treat it as internal.
50
+
51
+ ## Limits
52
+
53
+ The scan uses pattern matching, not a full Java parser, so counts can be slightly off for unusual code.
54
+ It finds tests by annotations such as `@Test`; Cucumber `.feature` files and Kotlin or Groovy sources are not read.
55
+
56
+ ## License
57
+
58
+ [MIT](LICENSE)
@@ -0,0 +1,48 @@
1
+ [build-system]
2
+ requires = ["setuptools>=77"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "sel2pw-scan"
7
+ version = "0.1.0"
8
+ description = "Offline inventory of Selenium Java test projects for a Playwright migration assessment"
9
+ readme = "README.md"
10
+ requires-python = ">=3.10"
11
+ license = "MIT"
12
+ license-files = ["LICENSE"]
13
+ authors = [{ name = "Shan Konduru" }]
14
+ keywords = ["selenium", "playwright", "migration", "assessment", "inventory", "test-automation", "java"]
15
+ classifiers = [
16
+ "Development Status :: 3 - Alpha",
17
+ "Environment :: Console",
18
+ "Intended Audience :: Developers",
19
+ "Operating System :: OS Independent",
20
+ "Programming Language :: Python :: 3",
21
+ "Programming Language :: Python :: 3 :: Only",
22
+ "Programming Language :: Python :: 3.10",
23
+ "Programming Language :: Python :: 3.11",
24
+ "Programming Language :: Python :: 3.12",
25
+ "Programming Language :: Python :: 3.13",
26
+ "Programming Language :: Python :: 3.14",
27
+ "Topic :: Software Development :: Quality Assurance",
28
+ "Topic :: Software Development :: Testing",
29
+ ]
30
+ dependencies = []
31
+
32
+ [project.optional-dependencies]
33
+ dev = ["pytest>=8"]
34
+
35
+ [project.urls]
36
+ Repository = "https://github.com/ShanKonduru/sel2pw-scan"
37
+ Issues = "https://github.com/ShanKonduru/sel2pw-scan/issues"
38
+ Changelog = "https://github.com/ShanKonduru/sel2pw-scan/blob/main/CHANGELOG.md"
39
+
40
+ [project.scripts]
41
+ sel2pw-scan = "sel2pw_scan.cli:main"
42
+
43
+ [tool.setuptools.packages.find]
44
+ where = ["src"]
45
+
46
+ [tool.pytest.ini_options]
47
+ testpaths = ["tests"]
48
+ norecursedirs = ["fixtures"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,3 @@
1
+ """sel2pw-scan: offline inventory of Selenium Java test projects for a Playwright migration assessment."""
2
+
3
+ __version__ = "0.1.0"
@@ -0,0 +1,3 @@
1
+ from .cli import main
2
+
3
+ raise SystemExit(main())
@@ -0,0 +1,46 @@
1
+ """Command line: ``sel2pw-scan SOURCE [--format md|json] [--out FILE]``."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import argparse
6
+ import sys
7
+ from pathlib import Path
8
+
9
+ from . import __version__
10
+ from .inventory import scan_project
11
+
12
+
13
+ def build_parser() -> argparse.ArgumentParser:
14
+ parser = argparse.ArgumentParser(
15
+ prog="sel2pw-scan",
16
+ description="Scan a Selenium Java test project and write an inventory report. "
17
+ "Runs offline: no AI model, no network calls.",
18
+ )
19
+ parser.add_argument("--version", action="version", version=f"sel2pw-scan {__version__}")
20
+ parser.add_argument("source", help="Folder of the Selenium Java project to scan")
21
+ parser.add_argument("--format", choices=["md", "json"], default="md",
22
+ help="md (default): summary report without code or locator values; "
23
+ "json: full detail, including locator values")
24
+ parser.add_argument("--out", help="Write the report to this file instead of printing it")
25
+ return parser
26
+
27
+
28
+ def main(argv: list[str] | None = None) -> int:
29
+ args = build_parser().parse_args(argv)
30
+ source = Path(args.source)
31
+ if not source.is_dir():
32
+ print(f"sel2pw-scan: not a folder: {source}", file=sys.stderr)
33
+ return 2
34
+ inv = scan_project(source)
35
+ text = inv.to_json() if args.format == "json" else inv.to_markdown()
36
+ if args.out:
37
+ Path(args.out).write_text(text, encoding="utf-8")
38
+ s = inv.summary()
39
+ print(f"wrote {args.out} ({s['java_files']} Java files, {s['test_methods']} test methods)")
40
+ else:
41
+ sys.stdout.write(text)
42
+ return 0
43
+
44
+
45
+ if __name__ == "__main__":
46
+ raise SystemExit(main())
@@ -0,0 +1,443 @@
1
+ """Deterministic inventory of a Selenium Java project.
2
+
3
+ This is a heuristic, regex-based scanner (no Java parser dependency). It gives
4
+ a grounded starting map for a migration assessment: test classes and methods,
5
+ lifecycle hooks, page objects, dependencies between project classes, and
6
+ Selenium constructs that need deliberate translation. Mapping IDs are
7
+ ``ClassName#method`` so they stay stable across runs.
8
+ """
9
+
10
+ from __future__ import annotations
11
+
12
+ import json
13
+ import re
14
+ from dataclasses import asdict, dataclass, field
15
+ from pathlib import Path
16
+
17
+ # Directories never scanned (build output, VCS and tool state).
18
+ SKIP_DIRS = {".sel2pw", ".git", ".venv", "venv", "node_modules", "target", "build", ".gradle", ".idea", "__pycache__", ".pytest_cache"}
19
+
20
+ # Selenium constructs that need deliberate translation when moving to Playwright.
21
+ CONSTRUCT_PATTERNS: dict[str, str] = {
22
+ "fixed_sleep": r"\bThread\.sleep\s*\(",
23
+ "explicit_wait": r"\bWebDriverWait\b|\bFluentWait\b",
24
+ "expected_conditions": r"\bExpectedConditions\.\w+",
25
+ "custom_expected_condition": r"\bExpectedCondition\s*<|\bnew\s+ExpectedCondition\b",
26
+ "implicit_wait": r"\bimplicitlyWait\s*\(",
27
+ "javascript_execution": r"\bJavascriptExecutor\b|\bexecuteScript\s*\(|\bexecuteAsyncScript\s*\(",
28
+ "frames": r"switchTo\(\)\s*\.\s*(?:frame|parentFrame|defaultContent)\s*\(",
29
+ "windows_tabs": r"getWindowHandles?\s*\(|switchTo\(\)\s*\.\s*(?:window|newWindow)\s*\(",
30
+ "alerts": r"switchTo\(\)\s*\.\s*alert\s*\(|\bAlert\b\s+\w+\s*=",
31
+ "actions_api": r"\bnew\s+Actions\s*\(",
32
+ "select_dropdown": r"\bnew\s+Select\s*\(",
33
+ "file_upload_or_download": r"\bLocalFileDetector\b|download\.default_directory|\.sendKeys\([^)]*(?:\.pdf|\.csv|\.xlsx?|\.png|\.jpe?g|\.txt|File\.separator|getAbsolutePath)",
34
+ "shadow_dom": r"getShadowRoot\s*\(|shadowRoot",
35
+ "screenshots": r"\bTakesScreenshot\b|getScreenshotAs\s*\(",
36
+ "cookies": r"manage\(\)\s*\.\s*(?:addCookie|getCookie|deleteCookie|deleteAllCookies)",
37
+ "browser_capabilities": r"\b(?:Chrome|Firefox|Edge|Safari)Options\b|\bDesiredCapabilities\b|\bMutableCapabilities\b",
38
+ "selenium_grid": r"\bRemoteWebDriver\b",
39
+ "page_factory": r"\bPageFactory\b|@FindBy\b|@FindBys\b|@FindAll\b",
40
+ "data_provider": r"@DataProvider\b|dataProvider\s*=|@ParameterizedTest\b|@(?:Csv|Value|Method|Enum)Source\b",
41
+ "listeners": r"@Listeners\b|\bITestListener\b|\bTestWatcher\b|@ExtendWith\b|\bWebDriverListener\b|\bEventFiringWebDriver\b",
42
+ "retries": r"retryAnalyzer\s*=|\bIRetryAnalyzer\b|@RepeatedTest\b|@RetryingTest\b",
43
+ "parallel_execution": r"parallel\s*=|threadPoolSize\s*=|@Execution\s*\(|\bThreadLocal\s*<\s*WebDriver",
44
+ "test_dependencies": r"dependsOnMethods\s*=|dependsOnGroups\s*=|@Order\s*\(|priority\s*=",
45
+ "soft_assertions": r"\bSoftAssert\b|\bassertAll\s*\(|\bSoftAssertions\b",
46
+ "hover_drag_drop": r"\.moveToElement\s*\(|\.dragAndDrop(?:By)?\s*\(|\.clickAndHold\s*\(",
47
+ "keyboard_keys": r"\bKeys\.\w+",
48
+ "stale_element_handling": r"StaleElementReferenceException",
49
+ }
50
+
51
+ LOCATOR_RE = re.compile(r'By\.(id|name|className|cssSelector|xpath|linkText|partialLinkText|tagName)\s*\(\s*"((?:[^"\\]|\\.)*)"')
52
+ FINDBY_RE = re.compile(r'@FindBy\s*\(\s*(id|name|className|css|xpath|linkText|partialLinkText|tagName|how)\s*=\s*(?:How\.\w+\s*,\s*using\s*=\s*)?"((?:[^"\\]|\\.)*)"')
53
+ PACKAGE_RE = re.compile(r"^\s*package\s+([\w.]+)\s*;", re.M)
54
+ IMPORT_RE = re.compile(r"^\s*import\s+(?:static\s+)?([\w.*]+)\s*;", re.M)
55
+ CLASS_RE = re.compile(r"\b(?:public\s+|abstract\s+|final\s+)*(class|interface|enum)\s+(\w+)(?:\s*<[^>{]*>)?(?:\s+extends\s+([\w.<>, ]+?))?(?:\s+implements\s+([\w.<>, ]+?))?\s*\{")
56
+ METHOD_RE = re.compile(
57
+ r"(?P<annos>(?:@\w+(?:\.\w+)*(?:\s*\((?:[^()]|\([^()]*\))*\))?\s*)*)"
58
+ r"(?:(?:public|protected|private|static|final|synchronized|abstract|default)\s+)*"
59
+ r"(?:<[^>]+>\s+)?(?P<ret>[\w.<>\[\],? ]+?)\s+(?P<name>\w+)\s*\((?P<params>[^()]*(?:\([^()]*\)[^()]*)*)\)\s*"
60
+ r"(?:throws\s+[\w.,\s]+)?\{"
61
+ )
62
+ ANNOTATION_RE = re.compile(r"@(\w+(?:\.\w+)*)(?:\s*\(((?:[^()]|\([^()]*\))*)\))?")
63
+
64
+ TEST_ANNOTATIONS = {"Test", "ParameterizedTest", "RepeatedTest", "TestFactory", "TestTemplate", "RetryingTest"}
65
+ LIFECYCLE_ANNOTATIONS = {
66
+ "Before", "After", "BeforeClass", "AfterClass", # JUnit 4
67
+ "BeforeEach", "AfterEach", "BeforeAll", "AfterAll", # JUnit 5
68
+ "BeforeMethod", "AfterMethod", "BeforeTest", "AfterTest", "BeforeSuite", "AfterSuite", "BeforeGroups", "AfterGroups", # TestNG
69
+ }
70
+ DISABLED_ANNOTATIONS = {"Ignore", "Disabled"}
71
+ KEYWORDS = {"if", "for", "while", "switch", "catch", "synchronized", "return", "new", "else", "try", "do"}
72
+
73
+
74
+ @dataclass
75
+ class JavaMethod:
76
+ id: str
77
+ name: str
78
+ kind: str # test | lifecycle | helper
79
+ line: int
80
+ annotations: list[str]
81
+ attributes: dict[str, str]
82
+ disabled: bool
83
+ constructs: list[str]
84
+ assertion_count: int
85
+ data_rows: int | None = None # data provider rows / parameterized cases, when statically countable
86
+
87
+
88
+ @dataclass
89
+ class JavaClass:
90
+ name: str
91
+ file: str
92
+ package: str
93
+ role: str # test | page_object | base_or_util | other
94
+ extends: str | None
95
+ framework: str | None
96
+ methods: list[JavaMethod] = field(default_factory=list)
97
+ constructs: dict[str, int] = field(default_factory=dict)
98
+ locators: list[dict[str, str]] = field(default_factory=list)
99
+ depends_on: list[str] = field(default_factory=list)
100
+
101
+
102
+ @dataclass
103
+ class Inventory:
104
+ root: str
105
+ java_files: int
106
+ frameworks: list[str]
107
+ selenium_version_hint: str | None
108
+ build_files: list[str]
109
+ classes: list[JavaClass]
110
+ construct_totals: dict[str, int]
111
+
112
+ @property
113
+ def tests(self) -> list[tuple[JavaClass, JavaMethod]]:
114
+ return [(c, m) for c in self.classes for m in c.methods if m.kind == "test"]
115
+
116
+ def to_dict(self) -> dict:
117
+ data = asdict(self)
118
+ data["summary"] = self.summary()
119
+ return data
120
+
121
+ def to_json(self) -> str:
122
+ return json.dumps(self.to_dict(), indent=2)
123
+
124
+ def summary(self) -> dict[str, int]:
125
+ roles: dict[str, int] = {}
126
+ for c in self.classes:
127
+ roles[c.role] = roles.get(c.role, 0) + 1
128
+ tests = self.tests
129
+ return {
130
+ "java_files": self.java_files,
131
+ "test_classes": roles.get("test", 0),
132
+ "page_objects": roles.get("page_object", 0),
133
+ "base_or_util_classes": roles.get("base_or_util", 0),
134
+ "test_methods": len(tests),
135
+ "disabled_test_methods": sum(1 for _, m in tests if m.disabled),
136
+ }
137
+
138
+ def to_markdown(self) -> str:
139
+ s = self.summary()
140
+ out = [
141
+ "# Selenium source inventory",
142
+ "",
143
+ f"- Root: `{self.root}`",
144
+ f"- Frameworks detected: {', '.join(self.frameworks) or 'unknown'}",
145
+ f"- Selenium version hint: {self.selenium_version_hint or 'not found in build files'}",
146
+ f"- Build files: {', '.join(self.build_files) or 'none'}",
147
+ f"- Java files: {s['java_files']} | test classes: {s['test_classes']} | page objects: {s['page_objects']} "
148
+ f"| base/util: {s['base_or_util_classes']} | test methods: {s['test_methods']} (disabled: {s['disabled_test_methods']})",
149
+ "",
150
+ "> Heuristic scan. Confirm against the source before relying on counts.",
151
+ "",
152
+ ]
153
+ if self.construct_totals:
154
+ out += ["## Constructs needing deliberate translation", "", "| Construct | Occurrences |", "|---|---|"]
155
+ out += [f"| {k} | {v} |" for k, v in sorted(self.construct_totals.items(), key=lambda kv: -kv[1])]
156
+ out.append("")
157
+ out += ["## Tests", "", "| ID | File | Line | Disabled | Assertions | Constructs |", "|---|---|---|---|---|---|"]
158
+ for cls, m in self.tests:
159
+ out.append(
160
+ f"| `{m.id}` | {cls.file} | {m.line} | {'yes' if m.disabled else ''} | {m.assertion_count} | {', '.join(m.constructs)} |"
161
+ )
162
+ out += ["", "## Classes", "", "| Class | Role | File | Extends | Depends on |", "|---|---|---|---|---|"]
163
+ for c in self.classes:
164
+ out.append(f"| {c.name} | {c.role} | {c.file} | {c.extends or ''} | {', '.join(c.depends_on)} |")
165
+ return "\n".join(out) + "\n"
166
+
167
+
168
+ # -- parsing helpers ---------------------------------------------------------
169
+
170
+ def mask_java(text: str) -> str:
171
+ """Blank out comments and string/char literal contents, keeping offsets."""
172
+ out = list(text)
173
+ i, n = 0, len(text)
174
+ while i < n:
175
+ c = text[i]
176
+ nxt = text[i + 1] if i + 1 < n else ""
177
+ if c == "/" and nxt == "/":
178
+ j = text.find("\n", i)
179
+ j = n if j == -1 else j
180
+ for k in range(i, j):
181
+ out[k] = " "
182
+ i = j
183
+ elif c == "/" and nxt == "*":
184
+ j = text.find("*/", i + 2)
185
+ j = n if j == -1 else j + 2
186
+ for k in range(i, j):
187
+ if out[k] != "\n":
188
+ out[k] = " "
189
+ i = j
190
+ elif c == '"' and text.startswith('"""', i):
191
+ j = text.find('"""', i + 3)
192
+ j = n if j == -1 else j + 3
193
+ for k in range(i + 3, max(j - 3, i + 3)):
194
+ if out[k] != "\n":
195
+ out[k] = " "
196
+ i = j
197
+ elif c in "\"'":
198
+ j = i + 1
199
+ while j < n and text[j] != c and text[j] != "\n":
200
+ j += 2 if text[j] == "\\" else 1
201
+ for k in range(i + 1, min(j, n)):
202
+ out[k] = " "
203
+ i = j + 1
204
+ else:
205
+ i += 1
206
+ return "".join(out)
207
+
208
+
209
+ def _match_brace(masked: str, open_idx: int) -> int:
210
+ depth = 0
211
+ for i in range(open_idx, len(masked)):
212
+ if masked[i] == "{":
213
+ depth += 1
214
+ elif masked[i] == "}":
215
+ depth -= 1
216
+ if depth == 0:
217
+ return i
218
+ return len(masked) - 1
219
+
220
+
221
+ _IMPORT_LINE_RE = re.compile(r"^\s*(?:import|package)\s.*$", re.M)
222
+
223
+
224
+ def _constructs(text: str) -> dict[str, int]:
225
+ text = _IMPORT_LINE_RE.sub("", text)
226
+ found = {}
227
+ for name, pattern in CONSTRUCT_PATTERNS.items():
228
+ count = len(re.findall(pattern, text))
229
+ if count:
230
+ found[name] = count
231
+ return found
232
+
233
+
234
+ def _detect_framework(imports: list[str]) -> str | None:
235
+ joined = " ".join(imports)
236
+ if "org.testng" in joined:
237
+ return "testng"
238
+ if "org.junit.jupiter" in joined:
239
+ return "junit5"
240
+ if "org.junit" in joined:
241
+ return "junit4"
242
+ return None
243
+
244
+
245
+ def _parse_annotations(block: str) -> tuple[list[str], dict[str, str]]:
246
+ names, attrs = [], {}
247
+ for m in ANNOTATION_RE.finditer(block):
248
+ name = m.group(1).split(".")[-1]
249
+ names.append(name)
250
+ if m.group(2):
251
+ for key, val in re.findall(r"(\w+)\s*=\s*(\"[^\"]*\"|\{[^}]*\}|[\w.]+)", m.group(2)):
252
+ attrs[f"{name}.{key}"] = val.strip('"')
253
+ if "=" not in m.group(2) and m.group(2).strip():
254
+ attrs[f"{name}.value"] = m.group(2).strip().strip('"')
255
+ return names, attrs
256
+
257
+
258
+ def _object_array_rows(masked_body: str) -> int | None:
259
+ """Rows in a ``new Object[][] { {...}, {...} }`` literal (TestNG data providers)."""
260
+ m = re.search(r"new\s+Object\s*\[\s*\]\s*\[\s*\]\s*\{", masked_body)
261
+ if not m:
262
+ return None
263
+ depth = rows = 0
264
+ for ch in masked_body[m.end() - 1:]:
265
+ if ch == "{":
266
+ depth += 1
267
+ if depth == 2:
268
+ rows += 1
269
+ elif ch == "}":
270
+ depth -= 1
271
+ if depth == 0:
272
+ return rows
273
+ return None
274
+
275
+
276
+ def _junit_param_rows(attrs: dict[str, str]) -> int | None:
277
+ for key in ("CsvSource.value", "ValueSource.strings", "ValueSource.ints", "ValueSource.longs", "ValueSource.doubles"):
278
+ if key in attrs:
279
+ value = attrs[key].strip()
280
+ quoted = re.findall(r'"(?:[^"\\]|\\.)*"', value)
281
+ if quoted:
282
+ return len(quoted)
283
+ inner = value.strip("{}").strip()
284
+ return len([v for v in inner.split(",") if v.strip()]) if inner else 0
285
+ return None
286
+
287
+
288
+ def _link_data_providers(methods: list[JavaMethod]) -> None:
289
+ providers = {}
290
+ for m in methods:
291
+ if "DataProvider" in m.annotations:
292
+ providers[m.attributes.get("DataProvider.name", m.name)] = m
293
+ for m in methods:
294
+ if m.kind != "test":
295
+ continue
296
+ name = m.attributes.get("Test.dataProvider")
297
+ if name and "Test.dataProviderClass" not in m.attributes and name in providers:
298
+ m.data_rows = providers[name].data_rows
299
+ elif m.data_rows is None:
300
+ m.data_rows = _junit_param_rows(m.attributes)
301
+
302
+
303
+ def _classify(cls_name: str, rel: str, methods: list[JavaMethod], body: str, extends: str | None) -> str:
304
+ if any(m.kind == "test" for m in methods):
305
+ return "test"
306
+ lower = cls_name.lower()
307
+ if "@FindBy" in body or lower.endswith(("page", "pageobject", "screen", "component", "section")) or "/pages/" in rel.lower():
308
+ return "page_object"
309
+ if lower.startswith("base") or lower.endswith(("base", "util", "utils", "helper", "helpers", "factory", "manager", "config", "listener", "driver")):
310
+ return "base_or_util"
311
+ if any(m.kind == "lifecycle" for m in methods):
312
+ return "base_or_util"
313
+ return "other"
314
+
315
+
316
+ def _scan_file(path: Path, rel: str) -> list[JavaClass]:
317
+ text = path.read_text(encoding="utf-8", errors="replace")
318
+ masked = mask_java(text)
319
+ package = (PACKAGE_RE.search(text) or [None, ""])[1]
320
+ imports = IMPORT_RE.findall(masked)
321
+ framework = _detect_framework(imports)
322
+ classes: list[JavaClass] = []
323
+
324
+ for cm in CLASS_RE.finditer(masked):
325
+ cls_name = cm.group(2)
326
+ open_idx = cm.end() - 1
327
+ close_idx = _match_brace(masked, open_idx)
328
+ body_masked = masked[open_idx:close_idx + 1]
329
+ body_text = text[open_idx:close_idx + 1]
330
+ # class-level annotations (e.g. @Listeners) sit just before the declaration
331
+ prefix_start = masked.rfind(";", 0, cm.start()) + 1
332
+ class_annos, class_attrs = _parse_annotations(masked[prefix_start:cm.start()])
333
+
334
+ methods: list[JavaMethod] = []
335
+ for mm in METHOD_RE.finditer(body_masked):
336
+ name = mm.group("name")
337
+ if name in KEYWORDS or name == cls_name:
338
+ continue
339
+ # only direct members: brace depth 1 relative to the class body
340
+ if body_masked[:mm.start()].count("{") - body_masked[:mm.start()].count("}") != 1:
341
+ continue
342
+ m_open = mm.end() - 1
343
+ m_close = _match_brace(body_masked, m_open)
344
+ # Include the annotation block so e.g. dataProvider= / retryAnalyzer= count for the method.
345
+ m_body = body_text[mm.start():m_close + 1]
346
+ # Annotation args are blanked in the masked text; read them from the original.
347
+ annos, attrs = _parse_annotations(body_text[mm.start():mm.start() + len(mm.group("annos") or "")])
348
+ if set(annos) & TEST_ANNOTATIONS:
349
+ kind = "test"
350
+ elif set(annos) & LIFECYCLE_ANNOTATIONS:
351
+ kind = "lifecycle"
352
+ else:
353
+ kind = "helper"
354
+ disabled = bool(set(annos) & DISABLED_ANNOTATIONS) or attrs.get("Test.enabled") == "false"
355
+ line = text.count("\n", 0, open_idx + mm.start("name")) + 1
356
+ methods.append(JavaMethod(
357
+ id=f"{cls_name}#{name}",
358
+ name=name,
359
+ kind=kind,
360
+ line=line,
361
+ annotations=annos,
362
+ attributes=attrs,
363
+ disabled=disabled,
364
+ constructs=sorted(_constructs(m_body)),
365
+ assertion_count=len(re.findall(r"\b(?:assert\w*|verify\w*|softly\.assert\w*|expect\w*)\s*\(", m_body)),
366
+ data_rows=_object_array_rows(body_masked[mm.start():m_close + 1]) if "DataProvider" in annos else None,
367
+ ))
368
+ _link_data_providers(methods)
369
+
370
+ if class_annos:
371
+ for m in methods:
372
+ if m.kind == "test" and ("Ignore" in class_annos or "Disabled" in class_annos):
373
+ m.disabled = True
374
+ extends = cm.group(3).strip() if cm.group(3) else None
375
+ locators = [{"strategy": s, "value": v} for s, v in LOCATOR_RE.findall(body_text)]
376
+ locators += [{"strategy": f"@FindBy.{s}", "value": v} for s, v in FINDBY_RE.findall(body_text)]
377
+ constructs = _constructs(body_text)
378
+ if class_annos:
379
+ constructs.update({k: v for k, v in _constructs(text[prefix_start:cm.start()]).items() if k not in constructs})
380
+ classes.append(JavaClass(
381
+ name=cls_name,
382
+ file=rel,
383
+ package=package,
384
+ role=_classify(cls_name, rel, methods, body_text, extends),
385
+ extends=extends,
386
+ framework=framework,
387
+ methods=methods,
388
+ constructs=constructs,
389
+ locators=locators,
390
+ ))
391
+ # Nested classes are found by the outer finditer too; that is fine for inventory purposes.
392
+ return classes
393
+
394
+
395
+ def _build_hints(root: Path) -> tuple[list[str], str | None]:
396
+ files, version = [], None
397
+ for name in ("pom.xml", "build.gradle", "build.gradle.kts", "testng.xml"):
398
+ for path in root.rglob(name):
399
+ if any(p in SKIP_DIRS for p in path.relative_to(root).parts[:-1]):
400
+ continue
401
+ files.append(path.relative_to(root).as_posix())
402
+ text = path.read_text(encoding="utf-8", errors="replace")
403
+ m = re.search(r"selenium-java</artifactId>\s*<version>([^<]+)</version>", text) or re.search(
404
+ r"selenium-java:([\w.\-${}]+)", text
405
+ ) or re.search(r"<selenium\.version>([^<]+)</selenium\.version>", text)
406
+ if m and not version:
407
+ version = m.group(1)
408
+ return sorted(files), version
409
+
410
+
411
+ def scan_project(root: str | Path) -> Inventory:
412
+ root = Path(root).resolve()
413
+ java_files = [
414
+ p for p in sorted(root.rglob("*.java"))
415
+ if not any(part in SKIP_DIRS for part in p.relative_to(root).parts[:-1])
416
+ ]
417
+ classes: list[JavaClass] = []
418
+ for path in java_files:
419
+ classes.extend(_scan_file(path, path.relative_to(root).as_posix()))
420
+
421
+ names = {c.name for c in classes}
422
+ for c in classes:
423
+ text = (root / c.file).read_text(encoding="utf-8", errors="replace")
424
+ masked = mask_java(text)
425
+ deps = {n for n in names if n != c.name and re.search(rf"\b{re.escape(n)}\b", masked)}
426
+ c.depends_on = sorted(deps)
427
+
428
+ # Per-file totals so nested classes are not counted twice.
429
+ totals: dict[str, int] = {}
430
+ for path in java_files:
431
+ for k, v in _constructs(path.read_text(encoding="utf-8", errors="replace")).items():
432
+ totals[k] = totals.get(k, 0) + v
433
+ frameworks = sorted({c.framework for c in classes if c.framework})
434
+ build_files, version = _build_hints(root)
435
+ return Inventory(
436
+ root=str(root),
437
+ java_files=len(java_files),
438
+ frameworks=frameworks,
439
+ selenium_version_hint=version,
440
+ build_files=build_files,
441
+ classes=classes,
442
+ construct_totals=totals,
443
+ )
@@ -0,0 +1,88 @@
1
+ Metadata-Version: 2.4
2
+ Name: sel2pw-scan
3
+ Version: 0.1.0
4
+ Summary: Offline inventory of Selenium Java test projects for a Playwright migration assessment
5
+ Author: Shan Konduru
6
+ License-Expression: MIT
7
+ Project-URL: Repository, https://github.com/ShanKonduru/sel2pw-scan
8
+ Project-URL: Issues, https://github.com/ShanKonduru/sel2pw-scan/issues
9
+ Project-URL: Changelog, https://github.com/ShanKonduru/sel2pw-scan/blob/main/CHANGELOG.md
10
+ Keywords: selenium,playwright,migration,assessment,inventory,test-automation,java
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Operating System :: OS Independent
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Programming Language :: Python :: 3 :: Only
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Programming Language :: Python :: 3.14
22
+ Classifier: Topic :: Software Development :: Quality Assurance
23
+ Classifier: Topic :: Software Development :: Testing
24
+ Requires-Python: >=3.10
25
+ Description-Content-Type: text/markdown
26
+ License-File: LICENSE
27
+ Provides-Extra: dev
28
+ Requires-Dist: pytest>=8; extra == "dev"
29
+ Dynamic: license-file
30
+
31
+ # sel2pw-scan
32
+
33
+ Offline inventory of a Selenium Java test project, for assessing a migration to Python Playwright.
34
+
35
+ `sel2pw-scan` reads your `.java` files and build files and writes a report: which test framework you use,
36
+ how many tests, page objects and helper classes you have, and how often your code uses Selenium features
37
+ that need deliberate work in a migration (fixed sleeps, frames, alerts, JavaScript calls, data providers and
38
+ so on).
39
+
40
+ It is a scanner only. It does not convert or change any code, uses no AI model and makes no network calls.
41
+
42
+ ## Install
43
+
44
+ Requires Python 3.10 or newer.
45
+
46
+ ```bash
47
+ pip install sel2pw-scan
48
+ ```
49
+
50
+ ## Use
51
+
52
+ ```bash
53
+ sel2pw-scan path/to/selenium-project --out inventory.md
54
+ ```
55
+
56
+ | Option | Effect |
57
+ |---|---|
58
+ | `--out FILE` | Write the report to a file instead of printing it |
59
+ | `--format md` | Default. Summary report, described below |
60
+ | `--format json` | Full detail for tooling, including locator values and annotation attributes |
61
+ | `--version` | Print the version |
62
+
63
+ `python -m sel2pw_scan ...` works the same way.
64
+
65
+ ## What the report contains
66
+
67
+ The Markdown report (`--format md`) lists:
68
+
69
+ - the scanned folder path, detected frameworks (TestNG, JUnit 4, JUnit 5), Selenium version and build files
70
+ - counts of Java files, test classes, page objects, base and utility classes, and test methods (including disabled ones)
71
+ - how often each of 27 Selenium constructs appears, such as `fixed_sleep`, `frames`, `alerts`,
72
+ `javascript_execution`, `data_provider` and `selenium_grid`
73
+ - every test method as `ClassName#method`, with its file, line, assertion count and constructs
74
+ - every class with its role, parent class and the project classes it depends on
75
+
76
+ It does **not** contain test code, locator values, string literals or test data. Read it before sharing, and
77
+ rename anything you consider sensitive, such as class names.
78
+
79
+ The JSON report (`--format json`) adds locator values and annotation attributes, so treat it as internal.
80
+
81
+ ## Limits
82
+
83
+ The scan uses pattern matching, not a full Java parser, so counts can be slightly off for unusual code.
84
+ It finds tests by annotations such as `@Test`; Cucumber `.feature` files and Kotlin or Groovy sources are not read.
85
+
86
+ ## License
87
+
88
+ [MIT](LICENSE)
@@ -0,0 +1,15 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ src/sel2pw_scan/__init__.py
5
+ src/sel2pw_scan/__main__.py
6
+ src/sel2pw_scan/cli.py
7
+ src/sel2pw_scan/inventory.py
8
+ src/sel2pw_scan.egg-info/PKG-INFO
9
+ src/sel2pw_scan.egg-info/SOURCES.txt
10
+ src/sel2pw_scan.egg-info/dependency_links.txt
11
+ src/sel2pw_scan.egg-info/entry_points.txt
12
+ src/sel2pw_scan.egg-info/requires.txt
13
+ src/sel2pw_scan.egg-info/top_level.txt
14
+ tests/test_cli.py
15
+ tests/test_inventory.py
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ sel2pw-scan = sel2pw_scan.cli:main
@@ -0,0 +1,3 @@
1
+
2
+ [dev]
3
+ pytest>=8
@@ -0,0 +1 @@
1
+ sel2pw_scan
@@ -0,0 +1,33 @@
1
+ import json
2
+ import subprocess
3
+ import sys
4
+
5
+ from sel2pw_scan import __version__
6
+ from sel2pw_scan.cli import main
7
+
8
+
9
+ def test_markdown_report_to_file(java_src, tmp_path, capsys):
10
+ out = tmp_path / "inventory.md"
11
+ assert main([str(java_src), "--out", str(out)]) == 0
12
+ assert "3 Java files, 3 test methods" in capsys.readouterr().out
13
+ md = out.read_text(encoding="utf-8")
14
+ assert md.startswith("# Selenium source inventory")
15
+ # The markdown report is the shareable one: no locator values.
16
+ assert "//button[text()='Sign in']" not in md
17
+
18
+
19
+ def test_json_report_to_stdout(java_src, capsys):
20
+ assert main([str(java_src), "--format", "json"]) == 0
21
+ data = json.loads(capsys.readouterr().out)
22
+ assert data["summary"]["test_methods"] == 3
23
+
24
+
25
+ def test_missing_folder_is_an_error(tmp_path, capsys):
26
+ assert main([str(tmp_path / "nope")]) == 2
27
+ assert "not a folder" in capsys.readouterr().err
28
+
29
+
30
+ def test_module_entry_point_and_version():
31
+ proc = subprocess.run([sys.executable, "-m", "sel2pw_scan", "--version"], capture_output=True, text=True)
32
+ assert proc.returncode == 0
33
+ assert proc.stdout.strip() == f"sel2pw-scan {__version__}"
@@ -0,0 +1,60 @@
1
+ import json
2
+
3
+ from sel2pw_scan.inventory import mask_java, scan_project
4
+
5
+
6
+ def test_scan_finds_tests_lifecycle_and_roles(java_src):
7
+ inv = scan_project(java_src)
8
+ s = inv.summary()
9
+ assert s["java_files"] == 3
10
+ assert s["test_classes"] == 1 and s["page_objects"] == 1 and s["base_or_util_classes"] == 1
11
+ assert [m.id for _, m in inv.tests] == [
12
+ "LoginTest#validLoginShowsDashboard",
13
+ "LoginTest#invalidLoginShowsError",
14
+ "LoginTest#logoutConfirmsWithAlert",
15
+ ]
16
+ assert inv.frameworks == ["testng"]
17
+ assert inv.selenium_version_hint == "4.21.0"
18
+
19
+ base = next(c for c in inv.classes if c.name == "BaseTest")
20
+ assert {m.name: m.kind for m in base.methods} == {"setUp": "lifecycle", "tearDown": "lifecycle"}
21
+ assert base.methods[1].attributes == {"AfterMethod.alwaysRun": "true"}
22
+
23
+
24
+ def test_method_details(java_src):
25
+ inv = scan_project(java_src)
26
+ tests = {m.name: m for _, m in inv.tests}
27
+ assert tests["logoutConfirmsWithAlert"].disabled is True
28
+ assert "alerts" in tests["logoutConfirmsWithAlert"].constructs
29
+ assert tests["invalidLoginShowsError"].attributes["Test.dataProvider"] == "badCredentials"
30
+ assert "data_provider" in tests["invalidLoginShowsError"].constructs
31
+ assert tests["validLoginShowsDashboard"].assertion_count == 2
32
+ assert "fixed_sleep" in tests["validLoginShowsDashboard"].constructs
33
+
34
+
35
+ def test_dependencies_and_locators(java_src):
36
+ inv = scan_project(java_src)
37
+ login_test = next(c for c in inv.classes if c.name == "LoginTest")
38
+ assert login_test.depends_on == ["BaseTest", "LoginPage"]
39
+ assert login_test.extends == "BaseTest"
40
+ page = next(c for c in inv.classes if c.name == "LoginPage")
41
+ values = {(l["strategy"], l["value"]) for l in page.locators}
42
+ assert ("xpath", "//button[text()='Sign in']") in values
43
+ assert ("@FindBy.css", "input[name='password']") in values
44
+ assert "frames" in page.constructs and "explicit_wait" in page.constructs
45
+
46
+
47
+ def test_json_and_markdown_render(java_src):
48
+ inv = scan_project(java_src)
49
+ data = json.loads(inv.to_json())
50
+ assert data["summary"]["test_methods"] == 3
51
+ md = inv.to_markdown()
52
+ assert "`LoginTest#invalidLoginShowsError`" in md
53
+
54
+
55
+ def test_mask_keeps_offsets_and_hides_strings_and_comments():
56
+ src = 'String s = "@Test {"; // @Test {\n/* @Test */ int x;'
57
+ masked = mask_java(src)
58
+ assert len(masked) == len(src)
59
+ assert "@Test" not in masked and "{" not in masked
60
+ assert masked.count("\n") == 1