codevariability 0.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- codevariability-0.2.0/CHANGELOG.md +27 -0
- codevariability-0.2.0/CONTRIBUTING.md +23 -0
- codevariability-0.2.0/LICENSE +21 -0
- codevariability-0.2.0/MANIFEST.in +10 -0
- codevariability-0.2.0/PKG-INFO +171 -0
- codevariability-0.2.0/PUBLICATION_CHECKLIST.md +29 -0
- codevariability-0.2.0/README.md +134 -0
- codevariability-0.2.0/docs/API.md +101 -0
- codevariability-0.2.0/docs/FORMATS.md +43 -0
- codevariability-0.2.0/docs/METRICS.md +68 -0
- codevariability-0.2.0/docs/RELEASING.md +50 -0
- codevariability-0.2.0/examples/basic.py +26 -0
- codevariability-0.2.0/examples/groups.py +31 -0
- codevariability-0.2.0/examples/inputs/variant_a.py +2 -0
- codevariability-0.2.0/examples/inputs/variant_b.py +2 -0
- codevariability-0.2.0/examples/interop.py +30 -0
- codevariability-0.2.0/pyproject.toml +72 -0
- codevariability-0.2.0/setup.cfg +4 -0
- codevariability-0.2.0/src/codevariability/__init__.py +23 -0
- codevariability-0.2.0/src/codevariability/analysis.py +544 -0
- codevariability-0.2.0/src/codevariability/ast_tree_edit.py +350 -0
- codevariability-0.2.0/src/codevariability/cli.py +75 -0
- codevariability-0.2.0/src/codevariability/exceptions.py +2 -0
- codevariability-0.2.0/src/codevariability/group_comparison.py +979 -0
- codevariability-0.2.0/src/codevariability/interop.py +94 -0
- codevariability-0.2.0/src/codevariability/metrics.py +89 -0
- codevariability-0.2.0/src/codevariability/normalization.py +257 -0
- codevariability-0.2.0/src/codevariability/output.py +76 -0
- codevariability-0.2.0/src/codevariability/spreadsheet.py +68 -0
- codevariability-0.2.0/src/codevariability/validation.py +115 -0
- codevariability-0.2.0/src/codevariability.egg-info/PKG-INFO +171 -0
- codevariability-0.2.0/src/codevariability.egg-info/SOURCES.txt +45 -0
- codevariability-0.2.0/src/codevariability.egg-info/dependency_links.txt +1 -0
- codevariability-0.2.0/src/codevariability.egg-info/entry_points.txt +2 -0
- codevariability-0.2.0/src/codevariability.egg-info/requires.txt +14 -0
- codevariability-0.2.0/src/codevariability.egg-info/top_level.txt +1 -0
- codevariability-0.2.0/tests/test_adversarial_properties.py +166 -0
- codevariability-0.2.0/tests/test_analysis.py +171 -0
- codevariability-0.2.0/tests/test_ast_tree_edit.py +248 -0
- codevariability-0.2.0/tests/test_export_security.py +81 -0
- codevariability-0.2.0/tests/test_group_comparison.py +247 -0
- codevariability-0.2.0/tests/test_input_validation.py +489 -0
- codevariability-0.2.0/tests/test_metric_reference.py +176 -0
- codevariability-0.2.0/tests/test_optional_adapter.py +12 -0
- codevariability-0.2.0/tests/test_release_regressions.py +111 -0
- codevariability-0.2.0/tests/test_tokenization.py +104 -0
- codevariability-0.2.0/tests/test_user_flow.py +268 -0
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.2.0 — 2026-10-03
|
|
4
|
+
|
|
5
|
+
- Correct empty Python AST input scores, including comment-only source and
|
|
6
|
+
empty Python Markdown fences. Advance Python AST normalization to v3 while
|
|
7
|
+
preserving the v2 TED formula and nonempty-tree behavior.
|
|
8
|
+
|
|
9
|
+
- Separate group means from statistic keys; expose `within_group_columns` for
|
|
10
|
+
collision-safe access and preserve legitimate group labels.
|
|
11
|
+
- Replace the generic quoted-token fallback with linear scanning; keep token
|
|
12
|
+
output semantics and existing normalization IDs.
|
|
13
|
+
- Reject optional JavaScript results when input hashes differ from the base
|
|
14
|
+
analysis. Correct source typing and add permanent adversarial/property tests.
|
|
15
|
+
- Require scikit-learn >=1.5.0 and Pygments >=2.20.0 at runtime, setuptools
|
|
16
|
+
>=83.0.0 for builds, and pytest >=9.0.3 for development to exclude known
|
|
17
|
+
affected upstream versions. Add a source type-check gate.
|
|
18
|
+
|
|
19
|
+
- Prepare the independent `codevariability` distribution with a `src` layout,
|
|
20
|
+
standalone tests, public examples, and explicit package contents.
|
|
21
|
+
- Preserve the Python API and metric definitions from the historical 0.1.2
|
|
22
|
+
source. Keep `import codevariability` and the `codevariability` command.
|
|
23
|
+
- Resolve the optional JavaScript adapter only through its installed command;
|
|
24
|
+
remove discovery of files outside the Python project.
|
|
25
|
+
- Document installation, outputs, resource limits, and migration from the
|
|
26
|
+
historical local source. This is the first release from the independent
|
|
27
|
+
repository.
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Use Python 3.10 or later. Create a virtual environment, install the development
|
|
4
|
+
extra, then run the independent Python checks:
|
|
5
|
+
|
|
6
|
+
```bash
|
|
7
|
+
python -m pip install -e '.[dev]'
|
|
8
|
+
python -m ruff check src tests examples
|
|
9
|
+
python -m pytest
|
|
10
|
+
python -m build
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
Node.js is optional. The JavaScript integration test skips when the separate
|
|
14
|
+
adapter command is unavailable; all Python tests and builds run independently.
|
|
15
|
+
|
|
16
|
+
For bug reports, include versions, selected metric names, and the smallest
|
|
17
|
+
reproducible inputs that can be shared publicly. Keep credentials and private
|
|
18
|
+
source code out of reports. Pull requests should describe observable behavior,
|
|
19
|
+
include relevant regression tests, and update the documentation.
|
|
20
|
+
|
|
21
|
+
Preserve metric identifiers when definitions remain equivalent. Changes to
|
|
22
|
+
normalization, parsing, or score definitions need a new identifier and a
|
|
23
|
+
migration note. See [release instructions](docs/RELEASING.md).
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Otávio Gomes
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
include LICENSE README.md CHANGELOG.md CONTRIBUTING.md PUBLICATION_CHECKLIST.md
|
|
2
|
+
recursive-include src *.py
|
|
3
|
+
recursive-include tests *.py
|
|
4
|
+
recursive-include examples *.py
|
|
5
|
+
recursive-include docs *.md
|
|
6
|
+
global-exclude *.py[cod]
|
|
7
|
+
global-exclude .DS_Store
|
|
8
|
+
prune .github
|
|
9
|
+
prune build
|
|
10
|
+
prune dist
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: codevariability
|
|
3
|
+
Version: 0.2.0
|
|
4
|
+
Summary: Source-code similarity matrices, representative rankings, and independent group comparison.
|
|
5
|
+
Author: Otávio Gomes
|
|
6
|
+
Maintainer: Otávio Gomes
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
Project-URL: Homepage, https://github.com/otaviouss/codevariability-py
|
|
9
|
+
Project-URL: Documentation, https://github.com/otaviouss/codevariability-py/tree/main/docs
|
|
10
|
+
Project-URL: Source, https://github.com/otaviouss/codevariability-py
|
|
11
|
+
Project-URL: Issues, https://github.com/otaviouss/codevariability-py/issues
|
|
12
|
+
Keywords: code similarity,source code analysis,ast,software metrics
|
|
13
|
+
Classifier: Development Status :: 3 - Alpha
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Topic :: Software Development
|
|
20
|
+
Requires-Python: >=3.10
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: numpy>=1.23
|
|
24
|
+
Requires-Dist: pandas>=2.0
|
|
25
|
+
Requires-Dist: scikit-learn>=1.5.0
|
|
26
|
+
Requires-Dist: rapidfuzz>=3.0
|
|
27
|
+
Requires-Dist: Pygments<3,>=2.20.0
|
|
28
|
+
Requires-Dist: openpyxl>=3.1
|
|
29
|
+
Provides-Extra: dev
|
|
30
|
+
Requires-Dist: pytest>=9.0.3; extra == "dev"
|
|
31
|
+
Requires-Dist: hypothesis>=6.168.3; extra == "dev"
|
|
32
|
+
Requires-Dist: mypy>=2.4; extra == "dev"
|
|
33
|
+
Requires-Dist: ruff>=0.14; extra == "dev"
|
|
34
|
+
Requires-Dist: build>=1.2; extra == "dev"
|
|
35
|
+
Requires-Dist: twine>=6; extra == "dev"
|
|
36
|
+
Dynamic: license-file
|
|
37
|
+
|
|
38
|
+
# CodeVariability for Python
|
|
39
|
+
|
|
40
|
+
CodeVariability compares UTF-8 source-code files and returns similarity
|
|
41
|
+
matrices, descriptive statistics, representative rankings, and comparisons
|
|
42
|
+
between independent groups. It helps you identify common structures and
|
|
43
|
+
unusual variants in a collection without executing the submitted code.
|
|
44
|
+
|
|
45
|
+
Version **0.2.0** is the first release of this independent repository. The API is alpha. The distribution name is `codevariability`;
|
|
46
|
+
the import name and command are `codevariability`.
|
|
47
|
+
|
|
48
|
+
## Installation
|
|
49
|
+
|
|
50
|
+
Python 3.10 or later is required. Install from PyPI:
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
python -m pip install codevariability==0.2.0
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
For development, install a checkout with `python -m pip install .`.
|
|
57
|
+
The historical 0.1.2 code is the implementation baseline; 0.2.0 is the first
|
|
58
|
+
public PyPI release. Use a new virtual environment when migrating from a
|
|
59
|
+
local historical installation.
|
|
60
|
+
|
|
61
|
+
## Quick start
|
|
62
|
+
|
|
63
|
+
This example creates its own inputs and works with an installed package:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
from pathlib import Path
|
|
67
|
+
from tempfile import TemporaryDirectory
|
|
68
|
+
from codevariability import analyze
|
|
69
|
+
|
|
70
|
+
with TemporaryDirectory() as directory:
|
|
71
|
+
folder = Path(directory)
|
|
72
|
+
(folder / "a.py").write_text("def total(a, b):\n return a + b\n", encoding="utf-8")
|
|
73
|
+
(folder / "b.py").write_text("def sum_values(x, y):\n return x + y\n", encoding="utf-8")
|
|
74
|
+
result = analyze(folder, metrics="all")
|
|
75
|
+
print(result.statistics)
|
|
76
|
+
print(result.ranking)
|
|
77
|
+
print(result.most_representative)
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
The runnable [example](https://github.com/otaviouss/codevariability-py/blob/main/examples/basic.py) also exports JSON, CSV, and Excel.
|
|
81
|
+
For the bundled input files, the CLI is:
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
codevariability analyze examples/inputs --metrics all --output analysis-output
|
|
85
|
+
codevariability analyze examples/inputs --metrics cosine jaccard --output analysis-output/report.xlsx
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
## Inputs and metrics
|
|
89
|
+
|
|
90
|
+
Pass a directory or a list of file paths to `analyze()`. Directory scans are
|
|
91
|
+
not recursive. File basenames must be unique. Files are decoded as UTF-8,
|
|
92
|
+
including an optional BOM. Use `extensions="py"` or a list of extensions to
|
|
93
|
+
filter a directory.
|
|
94
|
+
|
|
95
|
+
| Metric | What it compares |
|
|
96
|
+
| --- | --- |
|
|
97
|
+
| `cosine` | Word frequencies in the full document. |
|
|
98
|
+
| `jaccard` | Sets of words in the full document. |
|
|
99
|
+
| `lcs` | The longest common subsequence of code tokens. |
|
|
100
|
+
| `levenshtein` | Token sequences using unit edit costs. |
|
|
101
|
+
| `ast_tree_edit_similarity` | Ordered, normalized Python syntax trees. |
|
|
102
|
+
|
|
103
|
+
`metrics="all"` includes all applicable metrics. Python AST analysis requires
|
|
104
|
+
Python inputs; explicit selection on incompatible inputs raises `AnalysisError`.
|
|
105
|
+
Source files are read in full. For Markdown, token and AST metrics use the
|
|
106
|
+
identified fenced code blocks; textual metrics include the full document.
|
|
107
|
+
|
|
108
|
+
## Results and export
|
|
109
|
+
|
|
110
|
+
`AnalysisResult` exposes `matrices`, `statistics`, `representativeness`,
|
|
111
|
+
`ranking`, `rankings_by_metric`, `most_representative`, and `most_distinct`.
|
|
112
|
+
Similarities range from 0 to 1. Statistics use each unique pair once. Rankings
|
|
113
|
+
average within each available dimension and then give dimensions equal weight.
|
|
114
|
+
|
|
115
|
+
Use `result.export(directory, formats=("json", "csv"))` for machine-readable
|
|
116
|
+
results and tables, or `result.to_excel(path)` for a workbook. Outputs include
|
|
117
|
+
metric identifiers, normalization and runtime versions, and input hashes.
|
|
118
|
+
Spreadsheet exports escape formula-like labels; JSON retains the original
|
|
119
|
+
labels. See [input/output details](https://github.com/otaviouss/codevariability-py/blob/main/docs/FORMATS.md).
|
|
120
|
+
|
|
121
|
+
`compare_groups()` supports two groups or a mapping of two or more groups,
|
|
122
|
+
file-label permutation tests, and Holm-adjusted p-values. Each group needs at
|
|
123
|
+
least two files and groups must not share physical files. A complete example
|
|
124
|
+
is in [examples/groups.py](https://github.com/otaviouss/codevariability-py/blob/main/examples/groups.py).
|
|
125
|
+
|
|
126
|
+
## Optional JavaScript integration
|
|
127
|
+
|
|
128
|
+
The Python library works without Node.js. To request structural JavaScript or
|
|
129
|
+
TypeScript analysis through `compare_groups(..., include_ast=True)`, install
|
|
130
|
+
the independent [codevariability-js](https://github.com/otaviouss/codevariability-js)
|
|
131
|
+
package and make its command available on `PATH`. Alternatively, import a
|
|
132
|
+
single-metric JSON with `load_matrix_json()` and attach it with `with_metric()`.
|
|
133
|
+
The Python build and normal test suite require no JavaScript checkout.
|
|
134
|
+
|
|
135
|
+
## Limits and errors
|
|
136
|
+
|
|
137
|
+
Scores describe text, tokens, or syntax; they do not establish functional
|
|
138
|
+
equivalence, correctness, authorship, or copied-code percentages. AST
|
|
139
|
+
normalization removes concrete names and literal values. Different sources
|
|
140
|
+
can therefore have structural similarity 1.
|
|
141
|
+
|
|
142
|
+
Exact tree comparison can be expensive. The default `max_ted_cells=2_000_000`
|
|
143
|
+
limits tables for each non-identical AST pair; exceeding it raises an error,
|
|
144
|
+
without approximation. `None` removes the limit. It does not limit source
|
|
145
|
+
file size, the full pairwise matrix, or total CPU time. Use trusted output
|
|
146
|
+
directories and apply application-level resource limits for untrusted inputs.
|
|
147
|
+
|
|
148
|
+
Expected input and serialization errors raise `AnalysisError`. JavaScript
|
|
149
|
+
integration also supports `ast_timeout=120` seconds. See the
|
|
150
|
+
[API](https://github.com/otaviouss/codevariability-py/blob/main/docs/API.md) and [metric definitions](https://github.com/otaviouss/codevariability-py/blob/main/docs/METRICS.md) for details.
|
|
151
|
+
|
|
152
|
+
## Contributing and license
|
|
153
|
+
|
|
154
|
+
See [CONTRIBUTING.md](https://github.com/otaviouss/codevariability-py/blob/main/CONTRIBUTING.md) for development and
|
|
155
|
+
[PUBLICATION_CHECKLIST.md](https://github.com/otaviouss/codevariability-py/blob/main/PUBLICATION_CHECKLIST.md) for release preparation.
|
|
156
|
+
The project uses the [MIT license](https://github.com/otaviouss/codevariability-py/blob/main/LICENSE), copyright 2026 Otávio Gomes.
|
|
157
|
+
|
|
158
|
+
## Compatibility
|
|
159
|
+
|
|
160
|
+
Arbitrary distinct nonempty group names are
|
|
161
|
+
accepted; use `result.within_group_columns` to locate collision-safe within-group
|
|
162
|
+
means. Keep files unchanged while an analysis runs; optional JavaScript results
|
|
163
|
+
with different input hashes raise `AnalysisError`.
|
|
164
|
+
|
|
165
|
+
Runtime minimums are scikit-learn 1.5.0 and Pygments 2.20.0; build minimum is
|
|
166
|
+
setuptools 83.0.0 and development tests require pytest 9.0.3. These scopes are
|
|
167
|
+
separate. See [API](docs/API.md) and [metric limits](docs/METRICS.md).
|
|
168
|
+
|
|
169
|
+
Python AST normalization now uses v3 to identify corrected empty-program
|
|
170
|
+
behavior. The TED formula ID remains v2; use compatible normalization IDs when
|
|
171
|
+
comparing results across versions.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
# Python publication checklist
|
|
2
|
+
|
|
3
|
+
Prepared version: **0.2.0**. Distribution: **codevariability**.
|
|
4
|
+
Import/CLI: **codevariability**. A checked item records an executed check;
|
|
5
|
+
unchecked items are required before the first public release.
|
|
6
|
+
|
|
7
|
+
- [x] Public source and selected behavior tests reviewed.
|
|
8
|
+
- [x] README, API, metric, format, and release documentation reviewed.
|
|
9
|
+
- [x] MIT license and original copyright preserved.
|
|
10
|
+
- [x] Metadata, distribution name, import name, and single-source version checked.
|
|
11
|
+
- [x] Dependencies resolved and checked for known vulnerabilities.
|
|
12
|
+
- [x] Independent Python tests and lint passed.
|
|
13
|
+
- [ ] Run repository CI on the final remediated commit (local runtime matrix is validated separately).
|
|
14
|
+
- [x] Clean wheel and sdist builds completed.
|
|
15
|
+
- [x] Every wheel and sdist member inspected.
|
|
16
|
+
- [x] Wheel installed in a fresh environment; import and CLI checked.
|
|
17
|
+
- [x] Sdist rebuilt independently.
|
|
18
|
+
- [x] README code and all bundled examples executed.
|
|
19
|
+
- [x] New files checked for secrets, local paths, and internal materials.
|
|
20
|
+
- [x] No research inputs, manuscript files, caches, or generated results included.
|
|
21
|
+
- [x] New repository URLs verified with authenticated access.
|
|
22
|
+
- [x] PyPI project-name existence checked; recheck ownership/version before upload.
|
|
23
|
+
- [x] Make the GitHub repository public and verify links without authentication.
|
|
24
|
+
- [ ] Configure registry authentication or a trusted publisher for this new project.
|
|
25
|
+
- [ ] Choose the release date, publish, and verify installation from the registry.
|
|
26
|
+
|
|
27
|
+
The historical local implementation is the baseline; this distribution is
|
|
28
|
+
the first public PyPI release. See
|
|
29
|
+
[docs/RELEASING.md](docs/RELEASING.md).
|
|
@@ -0,0 +1,134 @@
|
|
|
1
|
+
# CodeVariability for Python
|
|
2
|
+
|
|
3
|
+
CodeVariability compares UTF-8 source-code files and returns similarity
|
|
4
|
+
matrices, descriptive statistics, representative rankings, and comparisons
|
|
5
|
+
between independent groups. It helps you identify common structures and
|
|
6
|
+
unusual variants in a collection without executing the submitted code.
|
|
7
|
+
|
|
8
|
+
Version **0.2.0** is the first release of this independent repository. The API is alpha. The distribution name is `codevariability`;
|
|
9
|
+
the import name and command are `codevariability`.
|
|
10
|
+
|
|
11
|
+
## Installation
|
|
12
|
+
|
|
13
|
+
Python 3.10 or later is required. Install from PyPI:
|
|
14
|
+
|
|
15
|
+
```bash
|
|
16
|
+
python -m pip install codevariability==0.2.0
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
For development, install a checkout with `python -m pip install .`.
|
|
20
|
+
The historical 0.1.2 code is the implementation baseline; 0.2.0 is the first
|
|
21
|
+
public PyPI release. Use a new virtual environment when migrating from a
|
|
22
|
+
local historical installation.
|
|
23
|
+
|
|
24
|
+
## Quick start
|
|
25
|
+
|
|
26
|
+
This example creates its own inputs and works with an installed package:
|
|
27
|
+
|
|
28
|
+
```python
|
|
29
|
+
from pathlib import Path
|
|
30
|
+
from tempfile import TemporaryDirectory
|
|
31
|
+
from codevariability import analyze
|
|
32
|
+
|
|
33
|
+
with TemporaryDirectory() as directory:
|
|
34
|
+
folder = Path(directory)
|
|
35
|
+
(folder / "a.py").write_text("def total(a, b):\n return a + b\n", encoding="utf-8")
|
|
36
|
+
(folder / "b.py").write_text("def sum_values(x, y):\n return x + y\n", encoding="utf-8")
|
|
37
|
+
result = analyze(folder, metrics="all")
|
|
38
|
+
print(result.statistics)
|
|
39
|
+
print(result.ranking)
|
|
40
|
+
print(result.most_representative)
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
The runnable [example](https://github.com/otaviouss/codevariability-py/blob/main/examples/basic.py) also exports JSON, CSV, and Excel.
|
|
44
|
+
For the bundled input files, the CLI is:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
codevariability analyze examples/inputs --metrics all --output analysis-output
|
|
48
|
+
codevariability analyze examples/inputs --metrics cosine jaccard --output analysis-output/report.xlsx
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
## Inputs and metrics
|
|
52
|
+
|
|
53
|
+
Pass a directory or a list of file paths to `analyze()`. Directory scans are
|
|
54
|
+
not recursive. File basenames must be unique. Files are decoded as UTF-8,
|
|
55
|
+
including an optional BOM. Use `extensions="py"` or a list of extensions to
|
|
56
|
+
filter a directory.
|
|
57
|
+
|
|
58
|
+
| Metric | What it compares |
|
|
59
|
+
| --- | --- |
|
|
60
|
+
| `cosine` | Word frequencies in the full document. |
|
|
61
|
+
| `jaccard` | Sets of words in the full document. |
|
|
62
|
+
| `lcs` | The longest common subsequence of code tokens. |
|
|
63
|
+
| `levenshtein` | Token sequences using unit edit costs. |
|
|
64
|
+
| `ast_tree_edit_similarity` | Ordered, normalized Python syntax trees. |
|
|
65
|
+
|
|
66
|
+
`metrics="all"` includes all applicable metrics. Python AST analysis requires
|
|
67
|
+
Python inputs; explicit selection on incompatible inputs raises `AnalysisError`.
|
|
68
|
+
Source files are read in full. For Markdown, token and AST metrics use the
|
|
69
|
+
identified fenced code blocks; textual metrics include the full document.
|
|
70
|
+
|
|
71
|
+
## Results and export
|
|
72
|
+
|
|
73
|
+
`AnalysisResult` exposes `matrices`, `statistics`, `representativeness`,
|
|
74
|
+
`ranking`, `rankings_by_metric`, `most_representative`, and `most_distinct`.
|
|
75
|
+
Similarities range from 0 to 1. Statistics use each unique pair once. Rankings
|
|
76
|
+
average within each available dimension and then give dimensions equal weight.
|
|
77
|
+
|
|
78
|
+
Use `result.export(directory, formats=("json", "csv"))` for machine-readable
|
|
79
|
+
results and tables, or `result.to_excel(path)` for a workbook. Outputs include
|
|
80
|
+
metric identifiers, normalization and runtime versions, and input hashes.
|
|
81
|
+
Spreadsheet exports escape formula-like labels; JSON retains the original
|
|
82
|
+
labels. See [input/output details](https://github.com/otaviouss/codevariability-py/blob/main/docs/FORMATS.md).
|
|
83
|
+
|
|
84
|
+
`compare_groups()` supports two groups or a mapping of two or more groups,
|
|
85
|
+
file-label permutation tests, and Holm-adjusted p-values. Each group needs at
|
|
86
|
+
least two files and groups must not share physical files. A complete example
|
|
87
|
+
is in [examples/groups.py](https://github.com/otaviouss/codevariability-py/blob/main/examples/groups.py).
|
|
88
|
+
|
|
89
|
+
## Optional JavaScript integration
|
|
90
|
+
|
|
91
|
+
The Python library works without Node.js. To request structural JavaScript or
|
|
92
|
+
TypeScript analysis through `compare_groups(..., include_ast=True)`, install
|
|
93
|
+
the independent [codevariability-js](https://github.com/otaviouss/codevariability-js)
|
|
94
|
+
package and make its command available on `PATH`. Alternatively, import a
|
|
95
|
+
single-metric JSON with `load_matrix_json()` and attach it with `with_metric()`.
|
|
96
|
+
The Python build and normal test suite require no JavaScript checkout.
|
|
97
|
+
|
|
98
|
+
## Limits and errors
|
|
99
|
+
|
|
100
|
+
Scores describe text, tokens, or syntax; they do not establish functional
|
|
101
|
+
equivalence, correctness, authorship, or copied-code percentages. AST
|
|
102
|
+
normalization removes concrete names and literal values. Different sources
|
|
103
|
+
can therefore have structural similarity 1.
|
|
104
|
+
|
|
105
|
+
Exact tree comparison can be expensive. The default `max_ted_cells=2_000_000`
|
|
106
|
+
limits tables for each non-identical AST pair; exceeding it raises an error,
|
|
107
|
+
without approximation. `None` removes the limit. It does not limit source
|
|
108
|
+
file size, the full pairwise matrix, or total CPU time. Use trusted output
|
|
109
|
+
directories and apply application-level resource limits for untrusted inputs.
|
|
110
|
+
|
|
111
|
+
Expected input and serialization errors raise `AnalysisError`. JavaScript
|
|
112
|
+
integration also supports `ast_timeout=120` seconds. See the
|
|
113
|
+
[API](https://github.com/otaviouss/codevariability-py/blob/main/docs/API.md) and [metric definitions](https://github.com/otaviouss/codevariability-py/blob/main/docs/METRICS.md) for details.
|
|
114
|
+
|
|
115
|
+
## Contributing and license
|
|
116
|
+
|
|
117
|
+
See [CONTRIBUTING.md](https://github.com/otaviouss/codevariability-py/blob/main/CONTRIBUTING.md) for development and
|
|
118
|
+
[PUBLICATION_CHECKLIST.md](https://github.com/otaviouss/codevariability-py/blob/main/PUBLICATION_CHECKLIST.md) for release preparation.
|
|
119
|
+
The project uses the [MIT license](https://github.com/otaviouss/codevariability-py/blob/main/LICENSE), copyright 2026 Otávio Gomes.
|
|
120
|
+
|
|
121
|
+
## Compatibility
|
|
122
|
+
|
|
123
|
+
Arbitrary distinct nonempty group names are
|
|
124
|
+
accepted; use `result.within_group_columns` to locate collision-safe within-group
|
|
125
|
+
means. Keep files unchanged while an analysis runs; optional JavaScript results
|
|
126
|
+
with different input hashes raise `AnalysisError`.
|
|
127
|
+
|
|
128
|
+
Runtime minimums are scikit-learn 1.5.0 and Pygments 2.20.0; build minimum is
|
|
129
|
+
setuptools 83.0.0 and development tests require pytest 9.0.3. These scopes are
|
|
130
|
+
separate. See [API](docs/API.md) and [metric limits](docs/METRICS.md).
|
|
131
|
+
|
|
132
|
+
Python AST normalization now uses v3 to identify corrected empty-program
|
|
133
|
+
behavior. The TED formula ID remains v2; use compatible normalization IDs when
|
|
134
|
+
comparing results across versions.
|
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
# Python API
|
|
2
|
+
|
|
3
|
+
Import the public API from `codevariability`. Submodules implement the metrics
|
|
4
|
+
and validation; they are not a separately versioned extension API.
|
|
5
|
+
|
|
6
|
+
## Analysis
|
|
7
|
+
|
|
8
|
+
```text
|
|
9
|
+
analyze(files_or_directory, metrics="all", *, extensions=None,
|
|
10
|
+
max_ted_cells=2_000_000) -> AnalysisResult
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
`files_or_directory` is a directory, an individual file, or a sequence of file
|
|
14
|
+
paths. `extensions` filters directory inputs only. `metrics` accepts one name,
|
|
15
|
+
a sequence, or `"all"`. Duplicate selected names are deduplicated.
|
|
16
|
+
|
|
17
|
+
For repeated use, `CodeDataset.from_directory(directory, extensions=None)`
|
|
18
|
+
and `CodeDataset.from_files(paths)` create an immutable mapping of labels to
|
|
19
|
+
paths. `dataset.analyze(metrics="all", max_ted_cells=2_000_000)` reads the
|
|
20
|
+
current file contents and returns a new result.
|
|
21
|
+
|
|
22
|
+
| AnalysisResult member | Return value or action |
|
|
23
|
+
| --- | --- |
|
|
24
|
+
| `files` / `metrics` | File labels / tuple of selected metric names. |
|
|
25
|
+
| `matrices` | Mapping from metric name to a square pandas DataFrame. |
|
|
26
|
+
| `statistics` | One row per metric, using unique off-diagonal pairs. |
|
|
27
|
+
| `representativeness` | Mean scores by metric and dimension, overall score, and row dispersion. |
|
|
28
|
+
| `ranking` | Overall ranking: `rank`, `file`, and `score`. |
|
|
29
|
+
| `rankings_by_metric` | One ranking DataFrame per metric. |
|
|
30
|
+
| `most_representative` / `most_distinct` | Dictionaries containing `file`, `score`, and `rank`. |
|
|
31
|
+
| `medoid(metric)` | The file with highest mean similarity for that metric. |
|
|
32
|
+
| `composite(weights="equal")` | A separate weighted similarity DataFrame; it does not replace the normal ranking. |
|
|
33
|
+
| `with_metric(name, matrix)` | A new result containing a validated external matrix. |
|
|
34
|
+
| `export(directory, formats=("json", "csv"))` | Write JSON and/or CSV files. |
|
|
35
|
+
| `to_excel(path)` | Write one workbook. |
|
|
36
|
+
|
|
37
|
+
Custom composite weights must cover exactly the selected metrics, be finite
|
|
38
|
+
and nonnegative, and have a positive sum. External matrices must have unique
|
|
39
|
+
matching labels, finite real entries in [0, 1], symmetry, and diagonal 1.
|
|
40
|
+
Boolean and string entries are rejected. Names and dimensions must not collide
|
|
41
|
+
with reserved result columns. Derived tables are recalculated from the current
|
|
42
|
+
matrices; augmented results copy their matrices and metadata.
|
|
43
|
+
|
|
44
|
+
## Comparing groups
|
|
45
|
+
|
|
46
|
+
```text
|
|
47
|
+
compare_groups(group_a, group_b=None, *, metrics="all",
|
|
48
|
+
group_names=("group_a", "group_b"), permutations=10_000,
|
|
49
|
+
random_state=42, include_ast=False, alpha=0.05,
|
|
50
|
+
progress=False, cache_dir=None, ast_timeout=120.0,
|
|
51
|
+
max_ted_cells=2_000_000)
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
Pass two inputs (paths or lists of files) for a `GroupComparisonResult`, or a
|
|
55
|
+
mapping from group name to input for a `MultiGroupComparisonResult`. The mapping
|
|
56
|
+
form requires at least two groups. Each group needs at least two files; groups
|
|
57
|
+
must not share files, including through hardlinks. `progress` accepts a boolean
|
|
58
|
+
or a callback receiving messages.
|
|
59
|
+
|
|
60
|
+
The two-group result exposes `summary`, `overview`, `interpretation`,
|
|
61
|
+
`by_metric`, `by_dimension`, `within_groups`, `between_groups`, and `p_values`.
|
|
62
|
+
The mapping result exposes `global_test`, `pairwise`, `within_groups`,
|
|
63
|
+
`between_groups`, `overview`, `pairwise_overview`, and `interpretation`.
|
|
64
|
+
Both support `print_report()`, `export(directory)`, and `to_excel(path)`.
|
|
65
|
+
|
|
66
|
+
Group names must be distinct, nonempty strings. They are labels, not statistic
|
|
67
|
+
keys. For two groups, `within_group_columns` returns the two summary columns
|
|
68
|
+
in group order and is also recorded in metadata. Usually these are
|
|
69
|
+
`within_<name>`. If a label would collide with a statistic or another label,
|
|
70
|
+
`__group` is appended until the column is unique. Thus a group named
|
|
71
|
+
`difference` remains valid without replacing the `within_difference` contrast.
|
|
72
|
+
Use this property when reading arbitrary user-supplied group names.
|
|
73
|
+
|
|
74
|
+
Tests permute file-level group labels with fixed group sizes and use the
|
|
75
|
+
Monte Carlo correction `(extremes + 1)/(permutations + 1)`. Two-group
|
|
76
|
+
homogeneity and separation tests are two-sided. Global multigroup homogeneity
|
|
77
|
+
uses the upper tail of dispersion among within-group means. Holm correction
|
|
78
|
+
is applied separately by hypothesis family; pairwise correction includes all
|
|
79
|
+
reported pairs and measures. Group weighting is equal. These tests assume
|
|
80
|
+
exchangeable, independent files; repeated or paired observations need another
|
|
81
|
+
permutation design. See [examples/groups.py](../examples/groups.py).
|
|
82
|
+
|
|
83
|
+
## JavaScript integration and errors
|
|
84
|
+
|
|
85
|
+
`include_ast=True` adds structural JavaScript/TypeScript analysis only when
|
|
86
|
+
Python AST analysis is not already present. The installed `codevariability-js`
|
|
87
|
+
command must be on PATH. No sibling directory or another checkout is searched.
|
|
88
|
+
`ast_timeout` must be a positive finite number of seconds; timeout or adapter
|
|
89
|
+
failure raises `AnalysisError`. `cache_dir` is passed to the optional adapter.
|
|
90
|
+
|
|
91
|
+
Keep inputs unchanged throughout group analysis, including progress callbacks.
|
|
92
|
+
The optional adapter's raw input hashes must match those used by the base
|
|
93
|
+
analysis; a changed snapshot raises `AnalysisError` rather than combining
|
|
94
|
+
results from different file contents.
|
|
95
|
+
|
|
96
|
+
`load_matrix_json(path)` returns `(metric_name, pandas.DataFrame)` for the
|
|
97
|
+
single-metric schema in [FORMATS.md](FORMATS.md). Attach it with `with_metric()`.
|
|
98
|
+
|
|
99
|
+
Expected invalid input, incompatible metric, parsing, resource-budget, or
|
|
100
|
+
serialization failures raise `AnalysisError`. Filesystem failures may also
|
|
101
|
+
raise `OSError`. The CLI reports expected failures with a nonzero exit status.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Inputs and outputs
|
|
2
|
+
|
|
3
|
+
Source files use UTF-8, optionally with BOM. Directory discovery is nonrecursive
|
|
4
|
+
and uses natural filename ordering with deterministic ties. All regular files
|
|
5
|
+
are considered unless an extension filter is supplied. Explicit file lists
|
|
6
|
+
must have unique basenames. No submitted program is executed.
|
|
7
|
+
|
|
8
|
+
## Analysis files
|
|
9
|
+
|
|
10
|
+
`export()` writes `analysis.json`, `statistics.csv`, `representativeness.csv`,
|
|
11
|
+
`ranking.csv`, `<metric>_similarity_matrix.csv`, and `<metric>_ranking.csv`.
|
|
12
|
+
Select only JSON or CSV with the `formats` parameter. JSON contains file names,
|
|
13
|
+
metrics, metadata, matrices, statistics, and rankings. Non-simple external
|
|
14
|
+
metric names receive a safe filename containing a hash suffix; the original
|
|
15
|
+
name is retained inside the result and JSON.
|
|
16
|
+
|
|
17
|
+
The workbook contains Summary, Metadata, Representativeness, Ranking,
|
|
18
|
+
Statistics, and one sheet per metric. Group exports have separate group
|
|
19
|
+
statistics files and an overview workbook.
|
|
20
|
+
|
|
21
|
+
File and group labels resembling formulas are escaped in CSV and Excel,
|
|
22
|
+
including row and column labels. Escaped CSV labels have a leading tab inside
|
|
23
|
+
quoted fields; workbook labels use a leading apostrophe. JSON keeps exact
|
|
24
|
+
labels and is the preferred interchange format when labels matter.
|
|
25
|
+
|
|
26
|
+
Individual output files are replaced atomically. Existing destination symlinks
|
|
27
|
+
are rejected. Multi-file export is not a transaction over the entire directory.
|
|
28
|
+
Output/cache directories should be trusted; the library does not confine all
|
|
29
|
+
filesystem operations to a sandbox root. Input hashes identify file bytes and
|
|
30
|
+
are included with runtime, metric, normalization, and aggregation metadata.
|
|
31
|
+
|
|
32
|
+
## External matrix JSON
|
|
33
|
+
|
|
34
|
+
The supported schema is `codevariability.matrix.v1`. It requires `metric`,
|
|
35
|
+
unique nonempty `files`, and a square real-number `matrix`. Optional `metric_id`
|
|
36
|
+
and `metadata` identify the adapter, normalization, and metric dimension.
|
|
37
|
+
The matrix must be symmetric, have diagonal 1, and contain finite values in
|
|
38
|
+
[0, 1]. Duplicate JSON keys, booleans, implicit numeric strings, and nonfinite
|
|
39
|
+
numbers are rejected. Multi-metric `codevariability.analysis.v1` output from
|
|
40
|
+
JavaScript is not accepted by `load_matrix_json()`; export one metric instead.
|
|
41
|
+
|
|
42
|
+
A self-contained loading and attachment example is in
|
|
43
|
+
[examples/interop.py](../examples/interop.py).
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# Metric definitions and limits
|
|
2
|
+
|
|
3
|
+
Text normalization applies Unicode NFC, case folding, and Unicode word
|
|
4
|
+
extraction to the entire document. `cosine` uses raw word-count vectors,
|
|
5
|
+
not TF-IDF; an empty vector has zero similarity to other files, even another
|
|
6
|
+
empty document. `jaccard` uses word sets, returning 1 for two empty sets.
|
|
7
|
+
All result matrices explicitly have diagonal 1.
|
|
8
|
+
|
|
9
|
+
For code-token metrics, `.py` and other source files are read in full.
|
|
10
|
+
Markdown uses identified fenced blocks; backtick and tilde fences are supported,
|
|
11
|
+
with the first word of the opening information string naming the language.
|
|
12
|
+
Unclosed fences raise an error. Markdown without code blocks has an empty code
|
|
13
|
+
sequence. Pygments handles known languages; unknown languages use a generic
|
|
14
|
+
deterministic tokenizer. Comments and whitespace are excluded.
|
|
15
|
+
|
|
16
|
+
`lcs` divides longest common subsequence length by the longer token sequence.
|
|
17
|
+
`levenshtein` uses unit insertion, deletion, and substitution costs and returns
|
|
18
|
+
`1 - distance / max_length`. Two empty sequences have similarity 1; an empty
|
|
19
|
+
and a nonempty sequence have similarity 0.
|
|
20
|
+
|
|
21
|
+
Python AST normalization preserves ordered syntax nodes, operators, literal
|
|
22
|
+
categories, and structural fields. Concrete identifiers, literal values,
|
|
23
|
+
comments, formatting, and locations are removed. Identified Python Markdown
|
|
24
|
+
fragments appear in order under a synthetic root. Empty or comment-only
|
|
25
|
+
fragments are ignored. Python grammar depends on the running interpreter.
|
|
26
|
+
|
|
27
|
+
`ast_tree_edit_similarity` uses exact ordered Zhang–Shasha tree edit distance
|
|
28
|
+
with unit insertion, deletion, and relabeling costs. For nonempty normalized
|
|
29
|
+
trees with `n` and `m` nodes:
|
|
30
|
+
|
|
31
|
+
```text
|
|
32
|
+
similarity = 1 - tree_edit_distance / (n + m - 1)
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
The denominator is a valid upper bound: remove non-root nodes, relabel the
|
|
36
|
+
root if necessary, and insert the target's non-root nodes. Empty/empty has
|
|
37
|
+
similarity 1; empty/nonempty has similarity 0. The score is symmetric and in
|
|
38
|
+
[0, 1], but does not establish semantic or functional equivalence.
|
|
39
|
+
|
|
40
|
+
Exact TED uses O(nm) table memory and shape-dependent computation. Its default
|
|
41
|
+
budget is 2,000,000 cells `(n+1)*(m+1)` for a non-identical pair. Equal trees
|
|
42
|
+
are checked first without those tables. `max_ted_cells=None` removes the
|
|
43
|
+
budget. Parsing, full matrices, file sizes, and permutations have separate
|
|
44
|
+
resource costs that the caller must manage.
|
|
45
|
+
|
|
46
|
+
Rankings first average metrics within each dimension (`textual`,
|
|
47
|
+
`syntactic_token_sequence`, `structural_ast`), then average available dimensions
|
|
48
|
+
equally. If an external AST node-frequency baseline and TED are both present,
|
|
49
|
+
only TED enters the structural dimension. `overall_dispersion` is the sample
|
|
50
|
+
standard deviation of a file's off-diagonal values in the equally weighted
|
|
51
|
+
dimension-average matrix; it is 0 when fewer than two values exist.
|
|
52
|
+
|
|
53
|
+
Internal IDs and normalization versions are recorded in metadata. Python AST
|
|
54
|
+
uses `ast_tree_edit_similarity_v2` and `python_normalized_ast_tree_v3`.
|
|
55
|
+
Compare results only when metric IDs, normalization, runtime/parser versions,
|
|
56
|
+
and selected dimensions are compatible.
|
|
57
|
+
|
|
58
|
+
The generic fallback scans quoted tokens in linear time with linear auxiliary
|
|
59
|
+
storage; token output semantics and `pygments_lexemes_with_generic_fallback_v2`
|
|
60
|
+
remain unchanged.
|
|
61
|
+
|
|
62
|
+
Version 0.2.0 advances Python AST normalization to v3 so that
|
|
63
|
+
empty/comment-only source and empty Python fences follow the documented scores:
|
|
64
|
+
empty/empty is 1 and empty/nonempty is 0. Nonempty trees and the TED formula
|
|
65
|
+
remain unchanged. The normalization helper retains its synthetic container;
|
|
66
|
+
the public metric treats a container without fragments as an empty tree.
|
|
67
|
+
This improvement does not make LCS, exact TED, permutation tests, or complete
|
|
68
|
+
pairwise matrices linear. Callers must still budget total work and input sizes.
|