codevariability 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. codevariability-0.2.0/CHANGELOG.md +27 -0
  2. codevariability-0.2.0/CONTRIBUTING.md +23 -0
  3. codevariability-0.2.0/LICENSE +21 -0
  4. codevariability-0.2.0/MANIFEST.in +10 -0
  5. codevariability-0.2.0/PKG-INFO +171 -0
  6. codevariability-0.2.0/PUBLICATION_CHECKLIST.md +29 -0
  7. codevariability-0.2.0/README.md +134 -0
  8. codevariability-0.2.0/docs/API.md +101 -0
  9. codevariability-0.2.0/docs/FORMATS.md +43 -0
  10. codevariability-0.2.0/docs/METRICS.md +68 -0
  11. codevariability-0.2.0/docs/RELEASING.md +50 -0
  12. codevariability-0.2.0/examples/basic.py +26 -0
  13. codevariability-0.2.0/examples/groups.py +31 -0
  14. codevariability-0.2.0/examples/inputs/variant_a.py +2 -0
  15. codevariability-0.2.0/examples/inputs/variant_b.py +2 -0
  16. codevariability-0.2.0/examples/interop.py +30 -0
  17. codevariability-0.2.0/pyproject.toml +72 -0
  18. codevariability-0.2.0/setup.cfg +4 -0
  19. codevariability-0.2.0/src/codevariability/__init__.py +23 -0
  20. codevariability-0.2.0/src/codevariability/analysis.py +544 -0
  21. codevariability-0.2.0/src/codevariability/ast_tree_edit.py +350 -0
  22. codevariability-0.2.0/src/codevariability/cli.py +75 -0
  23. codevariability-0.2.0/src/codevariability/exceptions.py +2 -0
  24. codevariability-0.2.0/src/codevariability/group_comparison.py +979 -0
  25. codevariability-0.2.0/src/codevariability/interop.py +94 -0
  26. codevariability-0.2.0/src/codevariability/metrics.py +89 -0
  27. codevariability-0.2.0/src/codevariability/normalization.py +257 -0
  28. codevariability-0.2.0/src/codevariability/output.py +76 -0
  29. codevariability-0.2.0/src/codevariability/spreadsheet.py +68 -0
  30. codevariability-0.2.0/src/codevariability/validation.py +115 -0
  31. codevariability-0.2.0/src/codevariability.egg-info/PKG-INFO +171 -0
  32. codevariability-0.2.0/src/codevariability.egg-info/SOURCES.txt +45 -0
  33. codevariability-0.2.0/src/codevariability.egg-info/dependency_links.txt +1 -0
  34. codevariability-0.2.0/src/codevariability.egg-info/entry_points.txt +2 -0
  35. codevariability-0.2.0/src/codevariability.egg-info/requires.txt +14 -0
  36. codevariability-0.2.0/src/codevariability.egg-info/top_level.txt +1 -0
  37. codevariability-0.2.0/tests/test_adversarial_properties.py +166 -0
  38. codevariability-0.2.0/tests/test_analysis.py +171 -0
  39. codevariability-0.2.0/tests/test_ast_tree_edit.py +248 -0
  40. codevariability-0.2.0/tests/test_export_security.py +81 -0
  41. codevariability-0.2.0/tests/test_group_comparison.py +247 -0
  42. codevariability-0.2.0/tests/test_input_validation.py +489 -0
  43. codevariability-0.2.0/tests/test_metric_reference.py +176 -0
  44. codevariability-0.2.0/tests/test_optional_adapter.py +12 -0
  45. codevariability-0.2.0/tests/test_release_regressions.py +111 -0
  46. codevariability-0.2.0/tests/test_tokenization.py +104 -0
  47. codevariability-0.2.0/tests/test_user_flow.py +268 -0
@@ -0,0 +1,27 @@
1
+ # Changelog
2
+
3
+ ## 0.2.0 — 2026-10-03
4
+
5
+ - Correct empty Python AST input scores, including comment-only source and
6
+ empty Python Markdown fences. Advance Python AST normalization to v3 while
7
+ preserving the v2 TED formula and nonempty-tree behavior.
8
+
9
+ - Separate group means from statistic keys; expose `within_group_columns` for
10
+ collision-safe access and preserve legitimate group labels.
11
+ - Replace the generic quoted-token fallback with linear scanning; keep token
12
+ output semantics and existing normalization IDs.
13
+ - Reject optional JavaScript results when input hashes differ from the base
14
+ analysis. Correct source typing and add permanent adversarial/property tests.
15
+ - Require scikit-learn >=1.5.0 and Pygments >=2.20.0 at runtime, setuptools
16
+ >=83.0.0 for builds, and pytest >=9.0.3 for development to exclude known
17
+ affected upstream versions. Add a source type-check gate.
18
+
19
+ - Prepare the independent `codevariability` distribution with a `src` layout,
20
+ standalone tests, public examples, and explicit package contents.
21
+ - Preserve the Python API and metric definitions from the historical 0.1.2
22
+ source. Keep `import codevariability` and the `codevariability` command.
23
+ - Resolve the optional JavaScript adapter only through its installed command;
24
+ remove discovery of files outside the Python project.
25
+ - Document installation, outputs, resource limits, and migration from the
26
+ historical local source. This is the first release from the independent
27
+ repository.
@@ -0,0 +1,23 @@
1
+ # Contributing
2
+
3
+ Use Python 3.10 or later. Create a virtual environment, install the development
4
+ extra, then run the independent Python checks:
5
+
6
+ ```bash
7
+ python -m pip install -e '.[dev]'
8
+ python -m ruff check src tests examples
9
+ python -m pytest
10
+ python -m build
11
+ ```
12
+
13
+ Node.js is optional. The JavaScript integration test skips when the separate
14
+ adapter command is unavailable; all Python tests and builds run independently.
15
+
16
+ For bug reports, include versions, selected metric names, and the smallest
17
+ reproducible inputs that can be shared publicly. Keep credentials and private
18
+ source code out of reports. Pull requests should describe observable behavior,
19
+ include relevant regression tests, and update the documentation.
20
+
21
+ Preserve metric identifiers when definitions remain equivalent. Changes to
22
+ normalization, parsing, or score definitions need a new identifier and a
23
+ migration note. See [release instructions](docs/RELEASING.md).
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Otávio Gomes
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,10 @@
1
+ include LICENSE README.md CHANGELOG.md CONTRIBUTING.md PUBLICATION_CHECKLIST.md
2
+ recursive-include src *.py
3
+ recursive-include tests *.py
4
+ recursive-include examples *.py
5
+ recursive-include docs *.md
6
+ global-exclude *.py[cod]
7
+ global-exclude .DS_Store
8
+ prune .github
9
+ prune build
10
+ prune dist
@@ -0,0 +1,171 @@
1
+ Metadata-Version: 2.4
2
+ Name: codevariability
3
+ Version: 0.2.0
4
+ Summary: Source-code similarity matrices, representative rankings, and independent group comparison.
5
+ Author: Otávio Gomes
6
+ Maintainer: Otávio Gomes
7
+ License-Expression: MIT
8
+ Project-URL: Homepage, https://github.com/otaviouss/codevariability-py
9
+ Project-URL: Documentation, https://github.com/otaviouss/codevariability-py/tree/main/docs
10
+ Project-URL: Source, https://github.com/otaviouss/codevariability-py
11
+ Project-URL: Issues, https://github.com/otaviouss/codevariability-py/issues
12
+ Keywords: code similarity,source code analysis,ast,software metrics
13
+ Classifier: Development Status :: 3 - Alpha
14
+ Classifier: Operating System :: OS Independent
15
+ Classifier: Programming Language :: Python :: 3 :: Only
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: Topic :: Software Development
20
+ Requires-Python: >=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: numpy>=1.23
24
+ Requires-Dist: pandas>=2.0
25
+ Requires-Dist: scikit-learn>=1.5.0
26
+ Requires-Dist: rapidfuzz>=3.0
27
+ Requires-Dist: Pygments<3,>=2.20.0
28
+ Requires-Dist: openpyxl>=3.1
29
+ Provides-Extra: dev
30
+ Requires-Dist: pytest>=9.0.3; extra == "dev"
31
+ Requires-Dist: hypothesis>=6.168.3; extra == "dev"
32
+ Requires-Dist: mypy>=2.4; extra == "dev"
33
+ Requires-Dist: ruff>=0.14; extra == "dev"
34
+ Requires-Dist: build>=1.2; extra == "dev"
35
+ Requires-Dist: twine>=6; extra == "dev"
36
+ Dynamic: license-file
37
+
38
+ # CodeVariability for Python
39
+
40
+ CodeVariability compares UTF-8 source-code files and returns similarity
41
+ matrices, descriptive statistics, representative rankings, and comparisons
42
+ between independent groups. It helps you identify common structures and
43
+ unusual variants in a collection without executing the submitted code.
44
+
45
+ Version **0.2.0** is the first release of this independent repository. The API is alpha. The distribution name is `codevariability`;
46
+ the import name and command are `codevariability`.
47
+
48
+ ## Installation
49
+
50
+ Python 3.10 or later is required. Install from PyPI:
51
+
52
+ ```bash
53
+ python -m pip install codevariability==0.2.0
54
+ ```
55
+
56
+ For development, install a checkout with `python -m pip install .`.
57
+ The historical 0.1.2 code is the implementation baseline; 0.2.0 is the first
58
+ public PyPI release. Use a new virtual environment when migrating from a
59
+ local historical installation.
60
+
61
+ ## Quick start
62
+
63
+ This example creates its own inputs and works with an installed package:
64
+
65
+ ```python
66
+ from pathlib import Path
67
+ from tempfile import TemporaryDirectory
68
+ from codevariability import analyze
69
+
70
+ with TemporaryDirectory() as directory:
71
+ folder = Path(directory)
72
+ (folder / "a.py").write_text("def total(a, b):\n return a + b\n", encoding="utf-8")
73
+ (folder / "b.py").write_text("def sum_values(x, y):\n return x + y\n", encoding="utf-8")
74
+ result = analyze(folder, metrics="all")
75
+ print(result.statistics)
76
+ print(result.ranking)
77
+ print(result.most_representative)
78
+ ```
79
+
80
+ The runnable [example](https://github.com/otaviouss/codevariability-py/blob/main/examples/basic.py) also exports JSON, CSV, and Excel.
81
+ For the bundled input files, the CLI is:
82
+
83
+ ```bash
84
+ codevariability analyze examples/inputs --metrics all --output analysis-output
85
+ codevariability analyze examples/inputs --metrics cosine jaccard --output analysis-output/report.xlsx
86
+ ```
87
+
88
+ ## Inputs and metrics
89
+
90
+ Pass a directory or a list of file paths to `analyze()`. Directory scans are
91
+ not recursive. File basenames must be unique. Files are decoded as UTF-8,
92
+ including an optional BOM. Use `extensions="py"` or a list of extensions to
93
+ filter a directory.
94
+
95
+ | Metric | What it compares |
96
+ | --- | --- |
97
+ | `cosine` | Word frequencies in the full document. |
98
+ | `jaccard` | Sets of words in the full document. |
99
+ | `lcs` | The longest common subsequence of code tokens. |
100
+ | `levenshtein` | Token sequences using unit edit costs. |
101
+ | `ast_tree_edit_similarity` | Ordered, normalized Python syntax trees. |
102
+
103
+ `metrics="all"` includes all applicable metrics. Python AST analysis requires
104
+ Python inputs; explicit selection on incompatible inputs raises `AnalysisError`.
105
+ Source files are read in full. For Markdown, token and AST metrics use the
106
+ identified fenced code blocks; textual metrics include the full document.
107
+
108
+ ## Results and export
109
+
110
+ `AnalysisResult` exposes `matrices`, `statistics`, `representativeness`,
111
+ `ranking`, `rankings_by_metric`, `most_representative`, and `most_distinct`.
112
+ Similarities range from 0 to 1. Statistics use each unique pair once. Rankings
113
+ average within each available dimension and then give dimensions equal weight.
114
+
115
+ Use `result.export(directory, formats=("json", "csv"))` for machine-readable
116
+ results and tables, or `result.to_excel(path)` for a workbook. Outputs include
117
+ metric identifiers, normalization and runtime versions, and input hashes.
118
+ Spreadsheet exports escape formula-like labels; JSON retains the original
119
+ labels. See [input/output details](https://github.com/otaviouss/codevariability-py/blob/main/docs/FORMATS.md).
120
+
121
+ `compare_groups()` supports two groups or a mapping of two or more groups,
122
+ file-label permutation tests, and Holm-adjusted p-values. Each group needs at
123
+ least two files and groups must not share physical files. A complete example
124
+ is in [examples/groups.py](https://github.com/otaviouss/codevariability-py/blob/main/examples/groups.py).
125
+
126
+ ## Optional JavaScript integration
127
+
128
+ The Python library works without Node.js. To request structural JavaScript or
129
+ TypeScript analysis through `compare_groups(..., include_ast=True)`, install
130
+ the independent [codevariability-js](https://github.com/otaviouss/codevariability-js)
131
+ package and make its command available on `PATH`. Alternatively, import a
132
+ single-metric JSON with `load_matrix_json()` and attach it with `with_metric()`.
133
+ The Python build and normal test suite require no JavaScript checkout.
134
+
135
+ ## Limits and errors
136
+
137
+ Scores describe text, tokens, or syntax; they do not establish functional
138
+ equivalence, correctness, authorship, or copied-code percentages. AST
139
+ normalization removes concrete names and literal values. Different sources
140
+ can therefore have structural similarity 1.
141
+
142
+ Exact tree comparison can be expensive. The default `max_ted_cells=2_000_000`
143
+ limits tables for each non-identical AST pair; exceeding it raises an error,
144
+ without approximation. `None` removes the limit. It does not limit source
145
+ file size, the full pairwise matrix, or total CPU time. Use trusted output
146
+ directories and apply application-level resource limits for untrusted inputs.
147
+
148
+ Expected input and serialization errors raise `AnalysisError`. JavaScript
149
+ integration also supports `ast_timeout=120` seconds. See the
150
+ [API](https://github.com/otaviouss/codevariability-py/blob/main/docs/API.md) and [metric definitions](https://github.com/otaviouss/codevariability-py/blob/main/docs/METRICS.md) for details.
151
+
152
+ ## Contributing and license
153
+
154
+ See [CONTRIBUTING.md](https://github.com/otaviouss/codevariability-py/blob/main/CONTRIBUTING.md) for development and
155
+ [PUBLICATION_CHECKLIST.md](https://github.com/otaviouss/codevariability-py/blob/main/PUBLICATION_CHECKLIST.md) for release preparation.
156
+ The project uses the [MIT license](https://github.com/otaviouss/codevariability-py/blob/main/LICENSE), copyright 2026 Otávio Gomes.
157
+
158
+ ## Compatibility
159
+
160
+ Arbitrary distinct nonempty group names are
161
+ accepted; use `result.within_group_columns` to locate collision-safe within-group
162
+ means. Keep files unchanged while an analysis runs; optional JavaScript results
163
+ with different input hashes raise `AnalysisError`.
164
+
165
+ Runtime minimums are scikit-learn 1.5.0 and Pygments 2.20.0; build minimum is
166
+ setuptools 83.0.0 and development tests require pytest 9.0.3. These scopes are
167
+ separate. See [API](docs/API.md) and [metric limits](docs/METRICS.md).
168
+
169
+ Python AST normalization now uses v3 to identify corrected empty-program
170
+ behavior. The TED formula ID remains v2; use compatible normalization IDs when
171
+ comparing results across versions.
@@ -0,0 +1,29 @@
1
+ # Python publication checklist
2
+
3
+ Prepared version: **0.2.0**. Distribution: **codevariability**.
4
+ Import/CLI: **codevariability**. A checked item records an executed check;
5
+ unchecked items are required before the first public release.
6
+
7
+ - [x] Public source and selected behavior tests reviewed.
8
+ - [x] README, API, metric, format, and release documentation reviewed.
9
+ - [x] MIT license and original copyright preserved.
10
+ - [x] Metadata, distribution name, import name, and single-source version checked.
11
+ - [x] Dependencies resolved and checked for known vulnerabilities.
12
+ - [x] Independent Python tests and lint passed.
13
+ - [ ] Run repository CI on the final remediated commit (local runtime matrix is validated separately).
14
+ - [x] Clean wheel and sdist builds completed.
15
+ - [x] Every wheel and sdist member inspected.
16
+ - [x] Wheel installed in a fresh environment; import and CLI checked.
17
+ - [x] Sdist rebuilt independently.
18
+ - [x] README code and all bundled examples executed.
19
+ - [x] New files checked for secrets, local paths, and internal materials.
20
+ - [x] No research inputs, manuscript files, caches, or generated results included.
21
+ - [x] New repository URLs verified with authenticated access.
22
+ - [x] PyPI project-name existence checked; recheck ownership/version before upload.
23
+ - [x] Make the GitHub repository public and verify links without authentication.
24
+ - [ ] Configure registry authentication or a trusted publisher for this new project.
25
+ - [ ] Choose the release date, publish, and verify installation from the registry.
26
+
27
+ The historical local implementation is the baseline; this distribution is
28
+ the first public PyPI release. See
29
+ [docs/RELEASING.md](docs/RELEASING.md).
@@ -0,0 +1,134 @@
1
+ # CodeVariability for Python
2
+
3
+ CodeVariability compares UTF-8 source-code files and returns similarity
4
+ matrices, descriptive statistics, representative rankings, and comparisons
5
+ between independent groups. It helps you identify common structures and
6
+ unusual variants in a collection without executing the submitted code.
7
+
8
+ Version **0.2.0** is the first release of this independent repository. The API is alpha. The distribution name is `codevariability`;
9
+ the import name and command are `codevariability`.
10
+
11
+ ## Installation
12
+
13
+ Python 3.10 or later is required. Install from PyPI:
14
+
15
+ ```bash
16
+ python -m pip install codevariability==0.2.0
17
+ ```
18
+
19
+ For development, install a checkout with `python -m pip install .`.
20
+ The historical 0.1.2 code is the implementation baseline; 0.2.0 is the first
21
+ public PyPI release. Use a new virtual environment when migrating from a
22
+ local historical installation.
23
+
24
+ ## Quick start
25
+
26
+ This example creates its own inputs and works with an installed package:
27
+
28
+ ```python
29
+ from pathlib import Path
30
+ from tempfile import TemporaryDirectory
31
+ from codevariability import analyze
32
+
33
+ with TemporaryDirectory() as directory:
34
+ folder = Path(directory)
35
+ (folder / "a.py").write_text("def total(a, b):\n return a + b\n", encoding="utf-8")
36
+ (folder / "b.py").write_text("def sum_values(x, y):\n return x + y\n", encoding="utf-8")
37
+ result = analyze(folder, metrics="all")
38
+ print(result.statistics)
39
+ print(result.ranking)
40
+ print(result.most_representative)
41
+ ```
42
+
43
+ The runnable [example](https://github.com/otaviouss/codevariability-py/blob/main/examples/basic.py) also exports JSON, CSV, and Excel.
44
+ For the bundled input files, the CLI is:
45
+
46
+ ```bash
47
+ codevariability analyze examples/inputs --metrics all --output analysis-output
48
+ codevariability analyze examples/inputs --metrics cosine jaccard --output analysis-output/report.xlsx
49
+ ```
50
+
51
+ ## Inputs and metrics
52
+
53
+ Pass a directory or a list of file paths to `analyze()`. Directory scans are
54
+ not recursive. File basenames must be unique. Files are decoded as UTF-8,
55
+ including an optional BOM. Use `extensions="py"` or a list of extensions to
56
+ filter a directory.
57
+
58
+ | Metric | What it compares |
59
+ | --- | --- |
60
+ | `cosine` | Word frequencies in the full document. |
61
+ | `jaccard` | Sets of words in the full document. |
62
+ | `lcs` | The longest common subsequence of code tokens. |
63
+ | `levenshtein` | Token sequences using unit edit costs. |
64
+ | `ast_tree_edit_similarity` | Ordered, normalized Python syntax trees. |
65
+
66
+ `metrics="all"` includes all applicable metrics. Python AST analysis requires
67
+ Python inputs; explicit selection on incompatible inputs raises `AnalysisError`.
68
+ Source files are read in full. For Markdown, token and AST metrics use the
69
+ identified fenced code blocks; textual metrics include the full document.
70
+
71
+ ## Results and export
72
+
73
+ `AnalysisResult` exposes `matrices`, `statistics`, `representativeness`,
74
+ `ranking`, `rankings_by_metric`, `most_representative`, and `most_distinct`.
75
+ Similarities range from 0 to 1. Statistics use each unique pair once. Rankings
76
+ average within each available dimension and then give dimensions equal weight.
77
+
78
+ Use `result.export(directory, formats=("json", "csv"))` for machine-readable
79
+ results and tables, or `result.to_excel(path)` for a workbook. Outputs include
80
+ metric identifiers, normalization and runtime versions, and input hashes.
81
+ Spreadsheet exports escape formula-like labels; JSON retains the original
82
+ labels. See [input/output details](https://github.com/otaviouss/codevariability-py/blob/main/docs/FORMATS.md).
83
+
84
+ `compare_groups()` supports two groups or a mapping of two or more groups,
85
+ file-label permutation tests, and Holm-adjusted p-values. Each group needs at
86
+ least two files and groups must not share physical files. A complete example
87
+ is in [examples/groups.py](https://github.com/otaviouss/codevariability-py/blob/main/examples/groups.py).
88
+
89
+ ## Optional JavaScript integration
90
+
91
+ The Python library works without Node.js. To request structural JavaScript or
92
+ TypeScript analysis through `compare_groups(..., include_ast=True)`, install
93
+ the independent [codevariability-js](https://github.com/otaviouss/codevariability-js)
94
+ package and make its command available on `PATH`. Alternatively, import a
95
+ single-metric JSON with `load_matrix_json()` and attach it with `with_metric()`.
96
+ The Python build and normal test suite require no JavaScript checkout.
97
+
98
+ ## Limits and errors
99
+
100
+ Scores describe text, tokens, or syntax; they do not establish functional
101
+ equivalence, correctness, authorship, or copied-code percentages. AST
102
+ normalization removes concrete names and literal values. Different sources
103
+ can therefore have structural similarity 1.
104
+
105
+ Exact tree comparison can be expensive. The default `max_ted_cells=2_000_000`
106
+ limits tables for each non-identical AST pair; exceeding it raises an error,
107
+ without approximation. `None` removes the limit. It does not limit source
108
+ file size, the full pairwise matrix, or total CPU time. Use trusted output
109
+ directories and apply application-level resource limits for untrusted inputs.
110
+
111
+ Expected input and serialization errors raise `AnalysisError`. JavaScript
112
+ integration also supports `ast_timeout=120` seconds. See the
113
+ [API](https://github.com/otaviouss/codevariability-py/blob/main/docs/API.md) and [metric definitions](https://github.com/otaviouss/codevariability-py/blob/main/docs/METRICS.md) for details.
114
+
115
+ ## Contributing and license
116
+
117
+ See [CONTRIBUTING.md](https://github.com/otaviouss/codevariability-py/blob/main/CONTRIBUTING.md) for development and
118
+ [PUBLICATION_CHECKLIST.md](https://github.com/otaviouss/codevariability-py/blob/main/PUBLICATION_CHECKLIST.md) for release preparation.
119
+ The project uses the [MIT license](https://github.com/otaviouss/codevariability-py/blob/main/LICENSE), copyright 2026 Otávio Gomes.
120
+
121
+ ## Compatibility
122
+
123
+ Arbitrary distinct nonempty group names are
124
+ accepted; use `result.within_group_columns` to locate collision-safe within-group
125
+ means. Keep files unchanged while an analysis runs; optional JavaScript results
126
+ with different input hashes raise `AnalysisError`.
127
+
128
+ Runtime minimums are scikit-learn 1.5.0 and Pygments 2.20.0; build minimum is
129
+ setuptools 83.0.0 and development tests require pytest 9.0.3. These scopes are
130
+ separate. See [API](docs/API.md) and [metric limits](docs/METRICS.md).
131
+
132
+ Python AST normalization now uses v3 to identify corrected empty-program
133
+ behavior. The TED formula ID remains v2; use compatible normalization IDs when
134
+ comparing results across versions.
@@ -0,0 +1,101 @@
1
+ # Python API
2
+
3
+ Import the public API from `codevariability`. Submodules implement the metrics
4
+ and validation; they are not a separately versioned extension API.
5
+
6
+ ## Analysis
7
+
8
+ ```text
9
+ analyze(files_or_directory, metrics="all", *, extensions=None,
10
+ max_ted_cells=2_000_000) -> AnalysisResult
11
+ ```
12
+
13
+ `files_or_directory` is a directory, an individual file, or a sequence of file
14
+ paths. `extensions` filters directory inputs only. `metrics` accepts one name,
15
+ a sequence, or `"all"`. Duplicate selected names are deduplicated.
16
+
17
+ For repeated use, `CodeDataset.from_directory(directory, extensions=None)`
18
+ and `CodeDataset.from_files(paths)` create an immutable mapping of labels to
19
+ paths. `dataset.analyze(metrics="all", max_ted_cells=2_000_000)` reads the
20
+ current file contents and returns a new result.
21
+
22
+ | AnalysisResult member | Return value or action |
23
+ | --- | --- |
24
+ | `files` / `metrics` | File labels / tuple of selected metric names. |
25
+ | `matrices` | Mapping from metric name to a square pandas DataFrame. |
26
+ | `statistics` | One row per metric, using unique off-diagonal pairs. |
27
+ | `representativeness` | Mean scores by metric and dimension, overall score, and row dispersion. |
28
+ | `ranking` | Overall ranking: `rank`, `file`, and `score`. |
29
+ | `rankings_by_metric` | One ranking DataFrame per metric. |
30
+ | `most_representative` / `most_distinct` | Dictionaries containing `file`, `score`, and `rank`. |
31
+ | `medoid(metric)` | The file with highest mean similarity for that metric. |
32
+ | `composite(weights="equal")` | A separate weighted similarity DataFrame; it does not replace the normal ranking. |
33
+ | `with_metric(name, matrix)` | A new result containing a validated external matrix. |
34
+ | `export(directory, formats=("json", "csv"))` | Write JSON and/or CSV files. |
35
+ | `to_excel(path)` | Write one workbook. |
36
+
37
+ Custom composite weights must cover exactly the selected metrics, be finite
38
+ and nonnegative, and have a positive sum. External matrices must have unique
39
+ matching labels, finite real entries in [0, 1], symmetry, and diagonal 1.
40
+ Boolean and string entries are rejected. Names and dimensions must not collide
41
+ with reserved result columns. Derived tables are recalculated from the current
42
+ matrices; augmented results copy their matrices and metadata.
43
+
44
+ ## Comparing groups
45
+
46
+ ```text
47
+ compare_groups(group_a, group_b=None, *, metrics="all",
48
+ group_names=("group_a", "group_b"), permutations=10_000,
49
+ random_state=42, include_ast=False, alpha=0.05,
50
+ progress=False, cache_dir=None, ast_timeout=120.0,
51
+ max_ted_cells=2_000_000)
52
+ ```
53
+
54
+ Pass two inputs (paths or lists of files) for a `GroupComparisonResult`, or a
55
+ mapping from group name to input for a `MultiGroupComparisonResult`. The mapping
56
+ form requires at least two groups. Each group needs at least two files; groups
57
+ must not share files, including through hardlinks. `progress` accepts a boolean
58
+ or a callback receiving messages.
59
+
60
+ The two-group result exposes `summary`, `overview`, `interpretation`,
61
+ `by_metric`, `by_dimension`, `within_groups`, `between_groups`, and `p_values`.
62
+ The mapping result exposes `global_test`, `pairwise`, `within_groups`,
63
+ `between_groups`, `overview`, `pairwise_overview`, and `interpretation`.
64
+ Both support `print_report()`, `export(directory)`, and `to_excel(path)`.
65
+
66
+ Group names must be distinct, nonempty strings. They are labels, not statistic
67
+ keys. For two groups, `within_group_columns` returns the two summary columns
68
+ in group order and is also recorded in metadata. Usually these are
69
+ `within_<name>`. If a label would collide with a statistic or another label,
70
+ `__group` is appended until the column is unique. Thus a group named
71
+ `difference` remains valid without replacing the `within_difference` contrast.
72
+ Use this property when reading arbitrary user-supplied group names.
73
+
74
+ Tests permute file-level group labels with fixed group sizes and use the
75
+ Monte Carlo correction `(extremes + 1)/(permutations + 1)`. Two-group
76
+ homogeneity and separation tests are two-sided. Global multigroup homogeneity
77
+ uses the upper tail of dispersion among within-group means. Holm correction
78
+ is applied separately by hypothesis family; pairwise correction includes all
79
+ reported pairs and measures. Group weighting is equal. These tests assume
80
+ exchangeable, independent files; repeated or paired observations need another
81
+ permutation design. See [examples/groups.py](../examples/groups.py).
82
+
83
+ ## JavaScript integration and errors
84
+
85
+ `include_ast=True` adds structural JavaScript/TypeScript analysis only when
86
+ Python AST analysis is not already present. The installed `codevariability-js`
87
+ command must be on PATH. No sibling directory or another checkout is searched.
88
+ `ast_timeout` must be a positive finite number of seconds; timeout or adapter
89
+ failure raises `AnalysisError`. `cache_dir` is passed to the optional adapter.
90
+
91
+ Keep inputs unchanged throughout group analysis, including progress callbacks.
92
+ The optional adapter's raw input hashes must match those used by the base
93
+ analysis; a changed snapshot raises `AnalysisError` rather than combining
94
+ results from different file contents.
95
+
96
+ `load_matrix_json(path)` returns `(metric_name, pandas.DataFrame)` for the
97
+ single-metric schema in [FORMATS.md](FORMATS.md). Attach it with `with_metric()`.
98
+
99
+ Expected invalid input, incompatible metric, parsing, resource-budget, or
100
+ serialization failures raise `AnalysisError`. Filesystem failures may also
101
+ raise `OSError`. The CLI reports expected failures with a nonzero exit status.
@@ -0,0 +1,43 @@
1
+ # Inputs and outputs
2
+
3
+ Source files use UTF-8, optionally with BOM. Directory discovery is nonrecursive
4
+ and uses natural filename ordering with deterministic ties. All regular files
5
+ are considered unless an extension filter is supplied. Explicit file lists
6
+ must have unique basenames. No submitted program is executed.
7
+
8
+ ## Analysis files
9
+
10
+ `export()` writes `analysis.json`, `statistics.csv`, `representativeness.csv`,
11
+ `ranking.csv`, `<metric>_similarity_matrix.csv`, and `<metric>_ranking.csv`.
12
+ Select only JSON or CSV with the `formats` parameter. JSON contains file names,
13
+ metrics, metadata, matrices, statistics, and rankings. Non-simple external
14
+ metric names receive a safe filename containing a hash suffix; the original
15
+ name is retained inside the result and JSON.
16
+
17
+ The workbook contains Summary, Metadata, Representativeness, Ranking,
18
+ Statistics, and one sheet per metric. Group exports have separate group
19
+ statistics files and an overview workbook.
20
+
21
+ File and group labels resembling formulas are escaped in CSV and Excel,
22
+ including row and column labels. Escaped CSV labels have a leading tab inside
23
+ quoted fields; workbook labels use a leading apostrophe. JSON keeps exact
24
+ labels and is the preferred interchange format when labels matter.
25
+
26
+ Individual output files are replaced atomically. Existing destination symlinks
27
+ are rejected. Multi-file export is not a transaction over the entire directory.
28
+ Output/cache directories should be trusted; the library does not confine all
29
+ filesystem operations to a sandbox root. Input hashes identify file bytes and
30
+ are included with runtime, metric, normalization, and aggregation metadata.
31
+
32
+ ## External matrix JSON
33
+
34
+ The supported schema is `codevariability.matrix.v1`. It requires `metric`,
35
+ unique nonempty `files`, and a square real-number `matrix`. Optional `metric_id`
36
+ and `metadata` identify the adapter, normalization, and metric dimension.
37
+ The matrix must be symmetric, have diagonal 1, and contain finite values in
38
+ [0, 1]. Duplicate JSON keys, booleans, implicit numeric strings, and nonfinite
39
+ numbers are rejected. Multi-metric `codevariability.analysis.v1` output from
40
+ JavaScript is not accepted by `load_matrix_json()`; export one metric instead.
41
+
42
+ A self-contained loading and attachment example is in
43
+ [examples/interop.py](../examples/interop.py).
@@ -0,0 +1,68 @@
1
+ # Metric definitions and limits
2
+
3
+ Text normalization applies Unicode NFC, case folding, and Unicode word
4
+ extraction to the entire document. `cosine` uses raw word-count vectors,
5
+ not TF-IDF; an empty vector has zero similarity to other files, even another
6
+ empty document. `jaccard` uses word sets, returning 1 for two empty sets.
7
+ All result matrices explicitly have diagonal 1.
8
+
9
+ For code-token metrics, `.py` and other source files are read in full.
10
+ Markdown uses identified fenced blocks; backtick and tilde fences are supported,
11
+ with the first word of the opening information string naming the language.
12
+ Unclosed fences raise an error. Markdown without code blocks has an empty code
13
+ sequence. Pygments handles known languages; unknown languages use a generic
14
+ deterministic tokenizer. Comments and whitespace are excluded.
15
+
16
+ `lcs` divides longest common subsequence length by the longer token sequence.
17
+ `levenshtein` uses unit insertion, deletion, and substitution costs and returns
18
+ `1 - distance / max_length`. Two empty sequences have similarity 1; an empty
19
+ and a nonempty sequence have similarity 0.
20
+
21
+ Python AST normalization preserves ordered syntax nodes, operators, literal
22
+ categories, and structural fields. Concrete identifiers, literal values,
23
+ comments, formatting, and locations are removed. Identified Python Markdown
24
+ fragments appear in order under a synthetic root. Empty or comment-only
25
+ fragments are ignored. Python grammar depends on the running interpreter.
26
+
27
+ `ast_tree_edit_similarity` uses exact ordered Zhang–Shasha tree edit distance
28
+ with unit insertion, deletion, and relabeling costs. For nonempty normalized
29
+ trees with `n` and `m` nodes:
30
+
31
+ ```text
32
+ similarity = 1 - tree_edit_distance / (n + m - 1)
33
+ ```
34
+
35
+ The denominator is a valid upper bound: remove non-root nodes, relabel the
36
+ root if necessary, and insert the target's non-root nodes. Empty/empty has
37
+ similarity 1; empty/nonempty has similarity 0. The score is symmetric and in
38
+ [0, 1], but does not establish semantic or functional equivalence.
39
+
40
+ Exact TED uses O(nm) table memory and shape-dependent computation. Its default
41
+ budget is 2,000,000 cells `(n+1)*(m+1)` for a non-identical pair. Equal trees
42
+ are checked first without those tables. `max_ted_cells=None` removes the
43
+ budget. Parsing, full matrices, file sizes, and permutations have separate
44
+ resource costs that the caller must manage.
45
+
46
+ Rankings first average metrics within each dimension (`textual`,
47
+ `syntactic_token_sequence`, `structural_ast`), then average available dimensions
48
+ equally. If an external AST node-frequency baseline and TED are both present,
49
+ only TED enters the structural dimension. `overall_dispersion` is the sample
50
+ standard deviation of a file's off-diagonal values in the equally weighted
51
+ dimension-average matrix; it is 0 when fewer than two values exist.
52
+
53
+ Internal IDs and normalization versions are recorded in metadata. Python AST
54
+ uses `ast_tree_edit_similarity_v2` and `python_normalized_ast_tree_v3`.
55
+ Compare results only when metric IDs, normalization, runtime/parser versions,
56
+ and selected dimensions are compatible.
57
+
58
+ The generic fallback scans quoted tokens in linear time with linear auxiliary
59
+ storage; token output semantics and `pygments_lexemes_with_generic_fallback_v2`
60
+ remain unchanged.
61
+
62
+ Version 0.2.0 advances Python AST normalization to v3 so that
63
+ empty/comment-only source and empty Python fences follow the documented scores:
64
+ empty/empty is 1 and empty/nonempty is 0. Nonempty trees and the TED formula
65
+ remain unchanged. The normalization helper retains its synthetic container;
66
+ the public metric treats a container without fragments as an empty tree.
67
+ This improvement does not make LCS, exact TED, permutation tests, or complete
68
+ pairwise matrices linear. Callers must still budget total work and input sizes.