compara 0.5.0__py3-none-win_amd64.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Binary file
@@ -0,0 +1,157 @@
1
+ Metadata-Version: 2.1
2
+ Name: compara
3
+ Version: 0.5.0
4
+ Summary: Compare two tables (CSV, NDJSON, Parquet) on a composite key: counts, per-column statistics and a self-contained HTML report.
5
+ Home-page: https://github.com/andrey-usa/compara
6
+ Project-URL: Source, https://github.com/andrey-usa/compara
7
+ Project-URL: Changelog, https://github.com/andrey-usa/compara/blob/HEAD/CHANGELOG.md
8
+ License: MIT
9
+ Keywords: csv,parquet,diff,compare,data-quality
10
+ Classifier: Programming Language :: Rust
11
+ Classifier: Environment :: Console
12
+ Classifier: License :: OSI Approved :: MIT License
13
+ Requires-Python: >=3.8
14
+ Description-Content-Type: text/markdown
15
+
16
+ # compara
17
+
18
+ Compare two tables on a composite key and get the counts, the per-column
19
+ statistics and a self-contained HTML report. Key columns, compared columns and
20
+ normalisation rules are parameters, so one tool serves every recurring
21
+ comparison: a nightly export against yesterday's, a CSV against the Parquet a
22
+ warehouse emits, a migration's source against its target.
23
+
24
+ - Reads **CSV, newline-delimited JSON and Parquet**, and the two sides need not
25
+ be in the same format.
26
+ - Ten million rows × 20 columns in about a second on four cores
27
+ ([numbers below](#speed)).
28
+ - One static binary; the report is a single HTML file with nothing to host.
29
+ - Exit codes made for pipelines: `0` identical, `1` differences, `2` error,
30
+ `3` duplicate keys (with `--fail-on-dups`).
31
+
32
+ ## Install
33
+
34
+ Prebuilt binaries for Linux (x86_64, aarch64), macOS (arm64, x86_64) and
35
+ Windows (x64) are on the [releases page](https://github.com/andrey-usa/compara/releases),
36
+ and the same binaries through package managers:
37
+
38
+ | | |
39
+ |---|---|
40
+ | Cargo | `cargo install --locked compara` |
41
+ | npm | `npm install -g compara` (or `npx compara …`) |
42
+ | pip / uv | `pip install compara` · `uv tool install compara` |
43
+ | Homebrew | `brew install andrey-usa/tap/compara` |
44
+ | Scoop | `scoop bucket add andrey-usa https://github.com/andrey-usa/scoop-bucket` then `scoop install compara` |
45
+
46
+ ## Use
47
+
48
+ ```sh
49
+ compara compare a.csv b.csv -k account_id,txn_id -i updated_at
50
+ compara compare a.csv b.csv -k account_id,txn_id --trim --tolerance 0.005
51
+ compara compare export.csv warehouse.parquet -k id --summary
52
+ compara compare a.csv b.csv --profile orders --json summary.json --export-dir out/
53
+ compara columns data.parquet # column names (reads only the footer)
54
+ compara head data.ndjson -n 5 # first rows, aligned
55
+ ```
56
+
57
+ `compare` writes `<a>__vs__<b>.html` (or `--out PATH`): counts, per-column
58
+ statistics, and the changed, added, removed and duplicate rows, each section
59
+ capped at `--max-rows` (counts are always exact). [`examples/report.html`](examples/report.html)
60
+ is one, made from the two files next to it:
61
+
62
+ ```sh
63
+ compara compare examples/orders_2026-08.csv examples/orders_2026-09.csv -k order_id,line_no -i updated_at
64
+ ```
65
+
66
+ | Option | Effect |
67
+ |---|---|
68
+ | `-k`, `--key` | composite key, required (or from a profile) |
69
+ | `-c`, `--compare` | columns to diff; default every column present in both files except the key |
70
+ | `-i`, `--ignore` | columns to skip (timestamps, run ids) |
71
+ | `--trim`, `--ignore-case`, `--empty-is-null` | normalisation before comparing (applies to key and values) |
72
+ | `--tolerance` | absolute numeric tolerance where both sides parse as numbers |
73
+ | `--delimiter`, `--encoding` | override auto-detection (CSV only) |
74
+ | `--threads` | how wide the default engine runs; the default is every core |
75
+ | `--memory` | memory to plan for (e.g. `8G`); decides when large files are joined in one pass instead of two |
76
+ | `--max-rows` | rows embedded per report section (default 50 000) |
77
+ | `--export-dir` | full, uncapped changed/added/removed CSVs |
78
+ | `--json` | also write a JSON summary (counts and column stats; schema in `docs/contract.schema.json`) |
79
+ | `--summary` | print the counts and write nothing |
80
+ | `--engine` | `auto` (default), `turbo`, `sortmerge` or `native` |
81
+ | `--no-compress` | plain JSON payload in the report, for pre-2023 browsers |
82
+ | `--fail-on-dups` | exit 3 when either file has duplicate keys |
83
+
84
+ Duplicate keys are counted and listed per file; the first occurrence of each
85
+ key takes part in the join.
86
+
87
+ **Profiles.** A `compara.toml` in the working directory (or
88
+ `~/.config/compara/compara.toml`, or `--config PATH`) names recurring
89
+ comparisons: key, compared and ignored columns, normalisation, engine. See
90
+ [`compara.example.toml`](compara.example.toml). The command line overrides the
91
+ profile.
92
+
93
+ ## Speed
94
+
95
+ Ten million rows × 20 columns, keyed on `(account_id, txn_id)`, `--ignore
96
+ updated_at`, one GitHub Actions runner (AMD EPYC 7763, 4 vCPU / 16 GB), best of
97
+ five ([run 36672430118](https://github.com/andrey-usa/data-comparison-polyglot/actions/runs/36672430118)).
98
+ Each cell is wall time · CPU time · memory above the mapped input.
99
+
100
+ | Format | Input a side | compara |
101
+ |---|---:|---|
102
+ | CSV | 3,509 MB | 1.20 s · 4.0 s · 591 MB |
103
+ | ndjson | 8,487 MB | 2.79 s · 10.4 s · 592 MB |
104
+ | Parquet | 2,074 MB | 1.41 s · 4.7 s · 1,342 MB |
105
+
106
+ compara started as the Rust port in
107
+ [data-comparison-polyglot](https://github.com/andrey-usa/data-comparison-polyglot),
108
+ where the same comparison is implemented in C, C++, Rust and Zig and measured
109
+ side by side; that repository has the full tables and how they are produced.
110
+
111
+ ## Engines
112
+
113
+ Three backends, one result contract: every engine returns identical counts and
114
+ column statistics for the same input, and the test suite asserts it cell for
115
+ cell. `--engine auto` takes the first one that can load the input, in this
116
+ order.
117
+
118
+ | Engine | Implementation | Memory model | Use it for |
119
+ |---|---|---|---|
120
+ | `turbo` | mapped file, byte-level index; CSV, JSON and Parquet | off-heap bytes, index in memory | the default, and the fastest |
121
+ | `sortmerge` | external sort, merge join | bounded memory, spills to disk | files past what memory can index |
122
+ | `native` | over the `csv` crate | in-memory, row-oriented | the dependency-light baseline |
123
+
124
+ All three read values as text, with no type inference, so `1.0` and `1` stay
125
+ different unless a tolerance is set, and all three treat an empty field as
126
+ absent whether or not it is quoted.
127
+
128
+ ### Input formats
129
+
130
+ `turbo` decides what a file is by its content rather than its name: a Parquet
131
+ magic number, a leading `{`, or a CSV header.
132
+
133
+ | Format | How it is read |
134
+ |---|---|
135
+ | CSV | mapped, scanned eight bytes at a time; a field is an offset and a length |
136
+ | newline-delimited JSON | the same, addressed by key rather than by column number; `\uXXXX` is decoded, so an escaped and a literal character compare equal |
137
+ | Parquet | pages decoded into an arena; dictionary-encoded columns point at the dictionary entry, so a repeated value costs eight bytes a row |
138
+
139
+ The Parquet reader is written here rather than taken from a crate, to keep that
140
+ representation instead of converting to Arrow arrays. It reads PLAIN,
141
+ dictionary (`PLAIN_DICTIONARY`, `RLE_DICTIONARY`), RLE booleans and the
142
+ version-2 delta encodings; uncompressed, snappy, gzip, zstd and LZ4-raw pages;
143
+ data pages v1 and v2; definition levels for optional columns. It refuses, by
144
+ name, nested or repeated schemas, `BYTE_STREAM_SPLIT`, LZO and Brotli.
145
+
146
+ A typed Parquet value is rendered to text by a fixed rule: byte arrays as their
147
+ bytes, booleans `true`/`false`, integers and decimals in full, floats in the
148
+ shortest round-trip form, dates `YYYY-MM-DD`, timestamps ISO 8601.
149
+ `tests/fixtures/formats/typed.csv` is that rule written down.
150
+
151
+ ## Contributing
152
+
153
+ See [CONTRIBUTING.md](CONTRIBUTING.md).
154
+
155
+ ## License
156
+
157
+ [MIT](LICENSE)
@@ -0,0 +1,5 @@
1
+ compara-0.5.0.data/scripts/compara.exe,sha256=AzTd77MJI-Oq6yOrTeiIHejXFVGu4JgG5V2MOMX0_Gk,1940480
2
+ compara-0.5.0.dist-info/METADATA,sha256=JUdUX5ffJ5LBqBdse3vQAWjw68WR05lIAOEfzXKA8LI,7499
3
+ compara-0.5.0.dist-info/WHEEL,sha256=g1aQ65onkFiWyh7QjiXnWXLnCV9xiZTddNDcljDxYOM,88
4
+ compara-0.5.0.dist-info/licenses/LICENSE,sha256=RSeVx-hjXtE643Nsk8gnOcomD9Wdf4PAXy7J7fSi4pQ,1067
5
+ compara-0.5.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: package.py
3
+ Root-Is-Purelib: false
4
+ Tag: py3-none-win_amd64
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 andrey-usa
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.