duplicatecode 0.1.0__py3-none-macosx_11_0_arm64.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,182 @@
1
+ Metadata-Version: 2.4
2
+ Name: duplicatecode
3
+ Version: 0.1.0
4
+ Classifier: Programming Language :: Rust
5
+ Classifier: Environment :: Console
6
+ Summary: Static (LLM-free) detection of duplicate/similar code in Python and TypeScript
7
+ License: MIT
8
+ Requires-Python: >=3.9
9
+ Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
10
+
11
+ # duplicatecode
12
+
13
+ Static (LLM-free) detection of duplicate / similar code in Python and TypeScript (TSX),
14
+ aimed at catching an LLM re-implementing something that already exists. Input can be a git diff
15
+ checked against existing source.
16
+
17
+ - `crates/duplicatecode-engine` – Rust engine (tree-sitter based)
18
+ - `crates/duplicatecode-cli` – `duplicatecode` CLI
19
+ - `dataset/` – benchmark: coding tasks + diverging implementations (see `dataset/README.md`)
20
+
21
+ ## Install
22
+
23
+ ```sh
24
+ uv tool install duplicatecode # from PyPI (prebuilt binary, no Rust needed)
25
+ uvx duplicatecode scan . # or run without installing
26
+ curl -fsSL https://raw.githubusercontent.com/bmsuisse/duplicatecode/main/install.sh | sh # standalone binary
27
+ ```
28
+
29
+ Binaries for Linux/macOS/Windows (x86_64, plus aarch64 on Linux/macOS) are also attached to each
30
+ [GitHub Release](https://github.com/bmsuisse/duplicatecode/releases). Releases are automatic: bump
31
+ `version` in `[workspace.package]` in `Cargo.toml`, merge to `main`, and CI tags and publishes it.
32
+
33
+ ## Usage
34
+
35
+ ```sh
36
+ cargo build --release
37
+ duplicatecode units src/ # list extracted units
38
+ git diff origin/main | duplicatecode diff --repo . # check added code against the repo
39
+ duplicatecode scan packages/ services/api # find duplicate groups inside one or more folders (monorepo)
40
+ duplicatecode scan . --exclude '**/generated/**' --skip-tests --fail-on-found # CI-style run
41
+ duplicatecode scan . --pairs --json # machine-readable pairs instead of groups
42
+ duplicatecode bench --dataset dataset [--file-level] [--mutations] [--negatives <other repo>]
43
+ ```
44
+
45
+ ## How it works
46
+
47
+ Units (functions, methods, classes, arrow-function components) are extracted with tree-sitter and
48
+ normalized (identifiers/literals abstracted; comments, docstrings and type annotations dropped).
49
+ Each pair gets several signals: token 4-/2-gram overlap, node-kind histogram, literals, API/attribute
50
+ names, name subwords, callees, and statement-level multiset/alignment scores. `combined` is a weighted
51
+ sum (0.6 structure, 0.2 name, 0.2 callees). Constructors and dunder methods are ignored by default.
52
+ An inverted index over k-grams keeps matching against large corpora fast.
53
+
54
+ ## Benchmark findings (Haiku vs Sonnet implementations of the same 50 tasks)
55
+
56
+ - Ranking is good (right match is top-1 in ~70–95% of cases) but absolute detection at a 1% false-positive
57
+ rate is modest for independently written code: ~35–50% (function level) of same-task pairs.
58
+ - Comparing whole files (which removes helper-splitting noise) gives 76–100% recall, mostly thanks to
59
+ names; structure-only signals reach ~35–55% on loose/hard tasks. Statement-level features are roughly on
60
+ par with token k-grams; a fitted logistic model does not beat the hand weights in cross-validation.
61
+ - Mechanical rewrites (rename, reorder, temp variable, logging, dead code) are handled well individually
62
+ (86–100% of pairs ≥ 0.7); all combined drops the mean score to ~0.58, mostly because renames zero the name signal.
63
+
64
+ ## Normalization
65
+
66
+ Before comparing, debug output (`print`, `logging.*`, `console.*`) is dropped; Python comprehensions are
67
+ read as the equivalent loop; `x += 1` as `x = x + 1`; a temp variable returned right after its definition
68
+ is inlined; operators are canonicalized (`===`/`==`/`is`, `and`/`&&`, `None`/`null`, `const`/`let`).
69
+ Mutation benchmark (`bench --mutations`): each of rename, statement swap, temp variable, logging, dead code and
70
+ loop-to-comprehension keeps >= 93% of pairs at score >= 0.7; all combined 40% (renames zero the name signal).
71
+
72
+ ## Real-repo validation (hand-judged)
73
+
74
+ Self-scans of three internal repositories (Python + TypeScript), 189 pairs judged by reading both units
75
+ (TRUE_DUP / PARTIAL / BOILERPLATE / FALSE):
76
+
77
+ - Raw scores are a weak signal: only ~23% of pairs above 0.55 were true duplicates, 40% incl. partial; below
78
+ 0.65 almost none. Real duplicates were almost all exact copies with matching names.
79
+ - Dominant false alarms: react-query/fetch wrappers differing by endpoint, per-entity CRUD endpoints, one-line
80
+ repository delegators, DTO/model classes, per-file test fixtures, generated clients.
81
+ - Filters added from this: field-only classes, constructors/dunders, generated files, tiny units (unless
82
+ near-identical with the same name), test code (unless identical), minimum name similarity (0.3).
83
+ - `--profile copies` (default) re-weights signals for copy-paste (name 0.36, statements 0.27, literals 0.18);
84
+ leave-one-repo-out AUC improved on all three repos (0.71/0.69/0.95 -> 0.75/0.79/0.96). At threshold 0.6 the
85
+ judged sample gives ~50% true duplicates and ~76% incl. partial while keeping ~70% of the true ones.
86
+ These numbers are partly in-sample (weights and filters were chosen on the same pairs) — re-judge a fresh sample
87
+ before trusting them. A structure-heavy `--profile reimpl` exists for renamed re-implementations
88
+ (use with `--min-name 0`) but has no real-repo precision data yet.
89
+
90
+ ### Fresh held-out check (copies profile, threshold 0.6)
91
+
92
+ A third, untouched sample of 57 pairs (nothing was tuned on it): 30% true duplicates, 47% incl. partial
93
+ (OneSales 0/24, MDMApp 11/24, CCMT2 6/9 true). The 50%/76% above was optimistic because it was measured on the
94
+ pairs used to choose the weights. Test code is about half of the noise but also holds real copies
95
+ (25% true either way), so it is kept by default; `--skip-tests` drops it.
96
+
97
+ ### Re-implementations of existing helpers (`reimpl-eval`)
98
+
99
+ 42 real helper functions from the three repos were described neutrally (no names/code) and re-implemented from scratch
100
+ by Haiku and Sonnet without seeing the repo. Fraction where the original is found (copies profile) and where the best
101
+ *other* match also crosses the threshold:
102
+
103
+ | threshold | Sonnet found | Haiku found | both | other-code alarm |
104
+ |---|---|---|---|---|
105
+ | 0.3 | 67% | 60% | 63% | 21% |
106
+ | 0.4 | 62% | 33% | 48% | 8% |
107
+ | 0.5 | 45% | 14% | 30% | 1% |
108
+ | 0.6 | 24% | 5% | 14% | 0% |
109
+
110
+ So `diff` (checking new code) defaults to threshold 0.4 while `scan` (existing copies) defaults to 0.6. About half of
111
+ independent re-implementations are caught; the rest are genuinely different code. The structure-heavy `reimpl`
112
+ profile is not better on this test.
113
+
114
+ ### Head-to-head: LLM-only vs LLM + CLI (MDMApp, OneSales)
115
+
116
+ Four Sonnet agents (report only, ~80 tool calls, max 30 groups) hunted duplicates in two repos: two with plain
117
+ read/grep, two with this CLI. All distinct groups (82) were then judged blind by independent agents that read the code.
118
+
119
+ | | groups | true dup. | precision (true / incl. partial) | pooled recall (true) | tool calls |
120
+ |---|---|---|---|---|---|
121
+ | MDMApp, LLM only | 25 | 15 | 60% / 80% | 60% | ~27–33 |
122
+ | MDMApp, with CLI | 26 | 20 | 77% / 92% | 80% | 7 |
123
+ | OneSales, LLM only | 28 | 10 | 36% / 71% | 67% | ~33–42 |
124
+ | OneSales, with CLI | 21 | 10 | 48% / 86% | 67% | 9 |
125
+
126
+ Only 18 of 82 groups were found by both, so the approaches are complementary. The agents' reports are capped at 30
127
+ groups, so the raw tool is a better measure of recall: `scan --threshold 0.6` (defaults) finds 38 of the 40 judged
128
+ true duplicates (25/25 MDMApp, 13/15 OneSales) and 23 of the 25 that the LLM-only agents found independently.
129
+ What it still misses: a differently-written picker function (same purpose, different code) and a formatFileSize
130
+ variant with different constants. Fixed after this test: tiny same-name exact copies (`min_tokens` 20 -> 8, near-exact
131
+ same-name rule) and multi-line module-level values such as `export const queryClient = new QueryClient({...})`.
132
+
133
+ ### Cheap model + CLI + skill (Haiku) vs Sonnet (MDMApp, OneSales)
134
+
135
+ Same task and judging as above (blind judges, pooled recall against all distinct confirmed true duplicates: 37 in MDMApp,
136
+ 18 in OneSales). `review` + `skills/duplicate-code-review/SKILL.md` were built from what the first runs missed.
137
+ "cost index" = tokens x relative price (Haiku assumed 1/3 of Sonnet per token; estimate).
138
+
139
+ | repo | approach | groups | precision true / incl. partial | recall | tokens | cost index |
140
+ |---|---|---|---|---|---|---|
141
+ | MDMApp | Sonnet, no CLI | 25 | 60% / 80% | 41% | 86k | 259 |
142
+ | MDMApp | Sonnet + CLI | 29 | 79% / 93% | 54% | 59k | 177 |
143
+ | MDMApp | **Haiku + CLI + skill** (r1 / r2) | 31 / 40 | 90% / 90% · 70% / 82% | 65% / **70%** | 86k / 86k | **86** |
144
+ | OneSales | Sonnet, no CLI | 28 | 36% / 71% | 56% | 131k | 393 |
145
+ | OneSales | Sonnet + CLI | 21 | 52% / 90% | 61% | 62k | 187 |
146
+ | OneSales | **Haiku + CLI + skill** (r1 / r2) | 11 / 30 | 82% / 91% · 33% / 67% | 50% / 50% | 88k / 74k | **74–88** |
147
+
148
+ r1 = first skill (top-60 groups only); r2 = skill with `--brief` paging. Read: on MDMApp Haiku + CLI beats Sonnet alone on
149
+ recall (70% vs 41%) and precision at about a third of the cost; on OneSales it is at par with Sonnet alone on recall
150
+ (50% vs 56%, one group) with better-or-equal precision at about a fifth of the cost, but below Sonnet + CLI. Broadening (r2)
151
+ buys recall at the price of precision on OneSales. Small samples; judges and finders are the same model family.
152
+
153
+ Round 3 (after `LIKELY-NOISE` tagging and the line-diff view; hard cap 30 groups): MDMApp 29 groups, precision 69% / 86%,
154
+ recall 51%, 69k tokens; OneSales 12 groups, precision 75% / 83%, recall 42%, 89k tokens. Recall pooled against 41 (MDMApp)
155
+ and 19 (OneSales) confirmed duplicates. Run-to-run variation of Haiku (59–63% / 51% on MDMApp, 42–47% on OneSales) is as
156
+ large as the differences between skill/tool versions, so the tag and diff view are not shown to help measurably; averaged over
157
+ three runs Haiku + CLI + skill lands at ~58% (MDMApp) and ~45% (OneSales) recall versus 37% / 53% for Sonnet alone and
158
+ 49% / 63% for Sonnet + CLI, at roughly a fifth to a third of the cost of Sonnet alone. Note the 30-group cap is itself a ceiling:
159
+ the confirmed pool holds 41 real duplicates in MDMApp.
160
+
161
+ ### Uncapped comparison (same 80-group cap, ~120-call budget for every arm) — final
162
+
163
+ Pooled confirmed true duplicates: 48 (MDMApp), 35 (OneSales); every group anyone reported was judged blind by reading the code.
164
+
165
+ | repo | arm | groups | precision true / incl. partial | recall | tokens | cost index (Haiku = 1/3 price) |
166
+ |---|---|---|---|---|---|---|
167
+ | MDMApp | Sonnet, no CLI | 65 | 42% / 65% | 65% | 174k | 522 |
168
+ | MDMApp | Sonnet + CLI + skill | 69 | 62% / 86% | 77% | 72k | 214 |
169
+ | MDMApp | Haiku + CLI + skill | 59 | 73% / 86% | 67% | 102k | **101** |
170
+ | OneSales | Sonnet, no CLI | 64 | 30% / 56% | 54% | 153k | 459 |
171
+ | OneSales | Sonnet + CLI + skill | 80 | 32% / 71% | 66% | 90k | 270 |
172
+ | OneSales | Haiku + CLI + skill | 47 | 38% / 68% | 40% | 85k | **84** |
173
+ | both | tool list only, IDENTICAL+NEAR-COPY, minus `LIKELY-NOISE` | 173 / 278 | (judged subset: 66% / 85% MDMApp, 38% / 64% OneSales) | **88% / 83%** | 0 | 0 |
174
+
175
+ - With the same budget Haiku + CLI matches Sonnet alone on MDMApp (67% vs 65% recall, better precision) at ~1/5 of the cost;
176
+ on OneSales it is below (40% vs 54%) at ~1/5.5 of the cost. Sonnet + CLI is best on recall in both.
177
+ - The unfiltered tool list already covers 83–88% of the pool: the agents' job is pruning, and their lower recall comes
178
+ from what they choose to report (they judge from brief lines and read few members), not from what the tool misses.
179
+ Precision of the raw list is unknown beyond the judged subset (which is biased towards groups an agent reported).
180
+ - Caveats: judges and finders are the same model family; the pool only contains duplicates someone found; single runs
181
+ per arm (Haiku varies by +-10 points between runs).
182
+
@@ -0,0 +1,5 @@
1
+ duplicatecode-0.1.0.data/scripts/duplicatecode,sha256=Q0cVPxA2r4C1dt0HZmDLHPquMjKWcv6wjA753ok2h2M,10072560
2
+ duplicatecode-0.1.0.dist-info/METADATA,sha256=a_PUyUeN14i_nGP5IyVuRbdpCbLCUsYM2JpM1FrfrAQ,11915
3
+ duplicatecode-0.1.0.dist-info/WHEEL,sha256=U927pSLvWHh5eaqcN5v72CdMm-HUWF344fKLe8HhSOU,102
4
+ duplicatecode-0.1.0.dist-info/sboms/duplicatecode.cyclonedx.json,sha256=1YMT8jTN7_lF_mij4sIBBB8TRjh76MrsL-crh5PIiqw,119902
5
+ duplicatecode-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: maturin (1.15.0)
3
+ Root-Is-Purelib: false
4
+ Tag: py3-none-macosx_11_0_arm64