compara 0.5.0__py3-none-win_amd64.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
Binary file
|
|
@@ -0,0 +1,157 @@
|
|
|
1
|
+
Metadata-Version: 2.1
|
|
2
|
+
Name: compara
|
|
3
|
+
Version: 0.5.0
|
|
4
|
+
Summary: Compare two tables (CSV, NDJSON, Parquet) on a composite key: counts, per-column statistics and a self-contained HTML report.
|
|
5
|
+
Home-page: https://github.com/andrey-usa/compara
|
|
6
|
+
Project-URL: Source, https://github.com/andrey-usa/compara
|
|
7
|
+
Project-URL: Changelog, https://github.com/andrey-usa/compara/blob/HEAD/CHANGELOG.md
|
|
8
|
+
License: MIT
|
|
9
|
+
Keywords: csv,parquet,diff,compare,data-quality
|
|
10
|
+
Classifier: Programming Language :: Rust
|
|
11
|
+
Classifier: Environment :: Console
|
|
12
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
13
|
+
Requires-Python: >=3.8
|
|
14
|
+
Description-Content-Type: text/markdown
|
|
15
|
+
|
|
16
|
+
# compara
|
|
17
|
+
|
|
18
|
+
Compare two tables on a composite key and get the counts, the per-column
|
|
19
|
+
statistics and a self-contained HTML report. Key columns, compared columns and
|
|
20
|
+
normalisation rules are parameters, so one tool serves every recurring
|
|
21
|
+
comparison: a nightly export against yesterday's, a CSV against the Parquet a
|
|
22
|
+
warehouse emits, a migration's source against its target.
|
|
23
|
+
|
|
24
|
+
- Reads **CSV, newline-delimited JSON and Parquet**, and the two sides need not
|
|
25
|
+
be in the same format.
|
|
26
|
+
- Ten million rows × 20 columns in about a second on four cores
|
|
27
|
+
([numbers below](#speed)).
|
|
28
|
+
- One static binary; the report is a single HTML file with nothing to host.
|
|
29
|
+
- Exit codes made for pipelines: `0` identical, `1` differences, `2` error,
|
|
30
|
+
`3` duplicate keys (with `--fail-on-dups`).
|
|
31
|
+
|
|
32
|
+
## Install
|
|
33
|
+
|
|
34
|
+
Prebuilt binaries for Linux (x86_64, aarch64), macOS (arm64, x86_64) and
|
|
35
|
+
Windows (x64) are on the [releases page](https://github.com/andrey-usa/compara/releases),
|
|
36
|
+
and the same binaries through package managers:
|
|
37
|
+
|
|
38
|
+
| | |
|
|
39
|
+
|---|---|
|
|
40
|
+
| Cargo | `cargo install --locked compara` |
|
|
41
|
+
| npm | `npm install -g compara` (or `npx compara …`) |
|
|
42
|
+
| pip / uv | `pip install compara` · `uv tool install compara` |
|
|
43
|
+
| Homebrew | `brew install andrey-usa/tap/compara` |
|
|
44
|
+
| Scoop | `scoop bucket add andrey-usa https://github.com/andrey-usa/scoop-bucket` then `scoop install compara` |
|
|
45
|
+
|
|
46
|
+
## Use
|
|
47
|
+
|
|
48
|
+
```sh
|
|
49
|
+
compara compare a.csv b.csv -k account_id,txn_id -i updated_at
|
|
50
|
+
compara compare a.csv b.csv -k account_id,txn_id --trim --tolerance 0.005
|
|
51
|
+
compara compare export.csv warehouse.parquet -k id --summary
|
|
52
|
+
compara compare a.csv b.csv --profile orders --json summary.json --export-dir out/
|
|
53
|
+
compara columns data.parquet # column names (reads only the footer)
|
|
54
|
+
compara head data.ndjson -n 5 # first rows, aligned
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
`compare` writes `<a>__vs__<b>.html` (or `--out PATH`): counts, per-column
|
|
58
|
+
statistics, and the changed, added, removed and duplicate rows, each section
|
|
59
|
+
capped at `--max-rows` (counts are always exact). [`examples/report.html`](examples/report.html)
|
|
60
|
+
is one, made from the two files next to it:
|
|
61
|
+
|
|
62
|
+
```sh
|
|
63
|
+
compara compare examples/orders_2026-08.csv examples/orders_2026-09.csv -k order_id,line_no -i updated_at
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
| Option | Effect |
|
|
67
|
+
|---|---|
|
|
68
|
+
| `-k`, `--key` | composite key, required (or from a profile) |
|
|
69
|
+
| `-c`, `--compare` | columns to diff; default every column present in both files except the key |
|
|
70
|
+
| `-i`, `--ignore` | columns to skip (timestamps, run ids) |
|
|
71
|
+
| `--trim`, `--ignore-case`, `--empty-is-null` | normalisation before comparing (applies to key and values) |
|
|
72
|
+
| `--tolerance` | absolute numeric tolerance where both sides parse as numbers |
|
|
73
|
+
| `--delimiter`, `--encoding` | override auto-detection (CSV only) |
|
|
74
|
+
| `--threads` | how wide the default engine runs; the default is every core |
|
|
75
|
+
| `--memory` | memory to plan for (e.g. `8G`); decides when large files are joined in one pass instead of two |
|
|
76
|
+
| `--max-rows` | rows embedded per report section (default 50 000) |
|
|
77
|
+
| `--export-dir` | full, uncapped changed/added/removed CSVs |
|
|
78
|
+
| `--json` | also write a JSON summary (counts and column stats; schema in `docs/contract.schema.json`) |
|
|
79
|
+
| `--summary` | print the counts and write nothing |
|
|
80
|
+
| `--engine` | `auto` (default), `turbo`, `sortmerge` or `native` |
|
|
81
|
+
| `--no-compress` | plain JSON payload in the report, for pre-2023 browsers |
|
|
82
|
+
| `--fail-on-dups` | exit 3 when either file has duplicate keys |
|
|
83
|
+
|
|
84
|
+
Duplicate keys are counted and listed per file; the first occurrence of each
|
|
85
|
+
key takes part in the join.
|
|
86
|
+
|
|
87
|
+
**Profiles.** A `compara.toml` in the working directory (or
|
|
88
|
+
`~/.config/compara/compara.toml`, or `--config PATH`) names recurring
|
|
89
|
+
comparisons: key, compared and ignored columns, normalisation, engine. See
|
|
90
|
+
[`compara.example.toml`](compara.example.toml). The command line overrides the
|
|
91
|
+
profile.
|
|
92
|
+
|
|
93
|
+
## Speed
|
|
94
|
+
|
|
95
|
+
Ten million rows × 20 columns, keyed on `(account_id, txn_id)`, `--ignore
|
|
96
|
+
updated_at`, one GitHub Actions runner (AMD EPYC 7763, 4 vCPU / 16 GB), best of
|
|
97
|
+
five ([run 36672430118](https://github.com/andrey-usa/data-comparison-polyglot/actions/runs/36672430118)).
|
|
98
|
+
Each cell is wall time · CPU time · memory above the mapped input.
|
|
99
|
+
|
|
100
|
+
| Format | Input a side | compara |
|
|
101
|
+
|---|---:|---|
|
|
102
|
+
| CSV | 3,509 MB | 1.20 s · 4.0 s · 591 MB |
|
|
103
|
+
| ndjson | 8,487 MB | 2.79 s · 10.4 s · 592 MB |
|
|
104
|
+
| Parquet | 2,074 MB | 1.41 s · 4.7 s · 1,342 MB |
|
|
105
|
+
|
|
106
|
+
compara started as the Rust port in
|
|
107
|
+
[data-comparison-polyglot](https://github.com/andrey-usa/data-comparison-polyglot),
|
|
108
|
+
where the same comparison is implemented in C, C++, Rust and Zig and measured
|
|
109
|
+
side by side; that repository has the full tables and how they are produced.
|
|
110
|
+
|
|
111
|
+
## Engines
|
|
112
|
+
|
|
113
|
+
Three backends, one result contract: every engine returns identical counts and
|
|
114
|
+
column statistics for the same input, and the test suite asserts it cell for
|
|
115
|
+
cell. `--engine auto` takes the first one that can load the input, in this
|
|
116
|
+
order.
|
|
117
|
+
|
|
118
|
+
| Engine | Implementation | Memory model | Use it for |
|
|
119
|
+
|---|---|---|---|
|
|
120
|
+
| `turbo` | mapped file, byte-level index; CSV, JSON and Parquet | off-heap bytes, index in memory | the default, and the fastest |
|
|
121
|
+
| `sortmerge` | external sort, merge join | bounded memory, spills to disk | files past what memory can index |
|
|
122
|
+
| `native` | over the `csv` crate | in-memory, row-oriented | the dependency-light baseline |
|
|
123
|
+
|
|
124
|
+
All three read values as text, with no type inference, so `1.0` and `1` stay
|
|
125
|
+
different unless a tolerance is set, and all three treat an empty field as
|
|
126
|
+
absent whether or not it is quoted.
|
|
127
|
+
|
|
128
|
+
### Input formats
|
|
129
|
+
|
|
130
|
+
`turbo` decides what a file is by its content rather than its name: a Parquet
|
|
131
|
+
magic number, a leading `{`, or a CSV header.
|
|
132
|
+
|
|
133
|
+
| Format | How it is read |
|
|
134
|
+
|---|---|
|
|
135
|
+
| CSV | mapped, scanned eight bytes at a time; a field is an offset and a length |
|
|
136
|
+
| newline-delimited JSON | the same, addressed by key rather than by column number; `\uXXXX` is decoded, so an escaped and a literal character compare equal |
|
|
137
|
+
| Parquet | pages decoded into an arena; dictionary-encoded columns point at the dictionary entry, so a repeated value costs eight bytes a row |
|
|
138
|
+
|
|
139
|
+
The Parquet reader is written here rather than taken from a crate, to keep that
|
|
140
|
+
representation instead of converting to Arrow arrays. It reads PLAIN,
|
|
141
|
+
dictionary (`PLAIN_DICTIONARY`, `RLE_DICTIONARY`), RLE booleans and the
|
|
142
|
+
version-2 delta encodings; uncompressed, snappy, gzip, zstd and LZ4-raw pages;
|
|
143
|
+
data pages v1 and v2; definition levels for optional columns. It refuses, by
|
|
144
|
+
name, nested or repeated schemas, `BYTE_STREAM_SPLIT`, LZO and Brotli.
|
|
145
|
+
|
|
146
|
+
A typed Parquet value is rendered to text by a fixed rule: byte arrays as their
|
|
147
|
+
bytes, booleans `true`/`false`, integers and decimals in full, floats in the
|
|
148
|
+
shortest round-trip form, dates `YYYY-MM-DD`, timestamps ISO 8601.
|
|
149
|
+
`tests/fixtures/formats/typed.csv` is that rule written down.
|
|
150
|
+
|
|
151
|
+
## Contributing
|
|
152
|
+
|
|
153
|
+
See [CONTRIBUTING.md](CONTRIBUTING.md).
|
|
154
|
+
|
|
155
|
+
## License
|
|
156
|
+
|
|
157
|
+
[MIT](LICENSE)
|
|
@@ -0,0 +1,5 @@
|
|
|
1
|
+
compara-0.5.0.data/scripts/compara.exe,sha256=AzTd77MJI-Oq6yOrTeiIHejXFVGu4JgG5V2MOMX0_Gk,1940480
|
|
2
|
+
compara-0.5.0.dist-info/METADATA,sha256=JUdUX5ffJ5LBqBdse3vQAWjw68WR05lIAOEfzXKA8LI,7499
|
|
3
|
+
compara-0.5.0.dist-info/WHEEL,sha256=g1aQ65onkFiWyh7QjiXnWXLnCV9xiZTddNDcljDxYOM,88
|
|
4
|
+
compara-0.5.0.dist-info/licenses/LICENSE,sha256=RSeVx-hjXtE643Nsk8gnOcomD9Wdf4PAXy7J7fSi4pQ,1067
|
|
5
|
+
compara-0.5.0.dist-info/RECORD,,
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 andrey-usa
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|