dference 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- dference-0.1.0/.gitignore +26 -0
- dference-0.1.0/CHANGELOG.md +61 -0
- dference-0.1.0/LICENSE +15 -0
- dference-0.1.0/PKG-INFO +283 -0
- dference-0.1.0/README.md +249 -0
- dference-0.1.0/examples/demo.py +599 -0
- dference-0.1.0/pyproject.toml +157 -0
- dference-0.1.0/src/dference/__init__.py +33 -0
- dference-0.1.0/src/dference/_compare.py +726 -0
- dference-0.1.0/src/dference/_query.py +492 -0
- dference-0.1.0/src/dference/_widget.py +251 -0
- dference-0.1.0/src/dference/py.typed +0 -0
- dference-0.1.0/src/dference/static/widget.css +952 -0
- dference-0.1.0/src/dference/static/widget.js +1370 -0
- dference-0.1.0/tests/conftest.py +38 -0
- dference-0.1.0/tests/e2e/__init__.py +0 -0
- dference-0.1.0/tests/e2e/conftest.py +277 -0
- dference-0.1.0/tests/e2e/test_ui.py +592 -0
- dference-0.1.0/tests/test_compare.py +351 -0
- dference-0.1.0/tests/test_properties.py +269 -0
- dference-0.1.0/tests/test_query.py +273 -0
- dference-0.1.0/tests/test_widget.py +146 -0
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
# Python
|
|
2
|
+
__pycache__/
|
|
3
|
+
*.py[cod]
|
|
4
|
+
.venv/
|
|
5
|
+
dist/
|
|
6
|
+
build/
|
|
7
|
+
*.egg-info/
|
|
8
|
+
|
|
9
|
+
# Frontend tooling
|
|
10
|
+
node_modules/
|
|
11
|
+
|
|
12
|
+
# Tooling caches
|
|
13
|
+
.ruff_cache/
|
|
14
|
+
.pytest_cache/
|
|
15
|
+
.coverage
|
|
16
|
+
coverage.xml
|
|
17
|
+
htmlcov/
|
|
18
|
+
|
|
19
|
+
# marimo
|
|
20
|
+
__marimo__/
|
|
21
|
+
.marimo/
|
|
22
|
+
|
|
23
|
+
# OS / editors
|
|
24
|
+
.DS_Store
|
|
25
|
+
.idea/
|
|
26
|
+
.vscode/
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the project uses
|
|
5
|
+
[Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
## [0.1.0] - 2026-10-06
|
|
10
|
+
|
|
11
|
+
First release.
|
|
12
|
+
|
|
13
|
+
### Added
|
|
14
|
+
|
|
15
|
+
- `compare()` – match two DataFrames on a (composite) key with a single polars
|
|
16
|
+
join and classify every row as equal, mismatch, only in left
|
|
17
|
+
(`missing_right`) or only in right (`missing_left`).
|
|
18
|
+
- `DiffResult` with `summary`, `columns`, `frame()`, `mismatches()` and
|
|
19
|
+
`column_stats()`.
|
|
20
|
+
- Exact comparison without tolerance; `null` and `NaN` count as equal missing
|
|
21
|
+
values.
|
|
22
|
+
- Strict dtypes: `strict=True` (the default) requires identical dtypes on both
|
|
23
|
+
sides; `strict=False` aligns lossless differences (integer width,
|
|
24
|
+
`Float32`/`Float64`, decimal precision, time units within one time zone, text
|
|
25
|
+
vs. `Categorical`/`Enum`, all-null columns). Anything else – including naive
|
|
26
|
+
vs. time-zone-aware and different time zones – raises one error listing each
|
|
27
|
+
column with a cast.
|
|
28
|
+
- Input support for polars (eager and lazy), pandas, pyarrow and objects with
|
|
29
|
+
the Arrow PyCapsule interface or `to_polars()`.
|
|
30
|
+
- `DataFrameDiff` anywidget for marimo and Jupyter. Filtering, sorting and
|
|
31
|
+
paging run in polars on the Python side; the browser only receives the
|
|
32
|
+
visible page. It has:
|
|
33
|
+
- a status overview on a shared scale, ordered left before right
|
|
34
|
+
- a status filter in the header of the status column; an active filter shows
|
|
35
|
+
as a removable pill in the toolbar
|
|
36
|
+
- quick filters `≠` and `=` per column (only rows that differ in, or are
|
|
37
|
+
equal in, that column) next to per-column match rates
|
|
38
|
+
- column menus with sorting and type-aware filters, where a row matches if
|
|
39
|
+
the left *or* the right value matches
|
|
40
|
+
- one colour and one chip (the side's short name) per side, used for
|
|
41
|
+
difference markers, statuses, one-sided columns, overview and detail view
|
|
42
|
+
- a side-by-side detail view with numeric deltas, stepping across pages
|
|
43
|
+
- invisible-character highlighting
|
|
44
|
+
- search, CSV export of the filtered rows, row selection synced to Python
|
|
45
|
+
(`selected_ids`, `selected_frame()`) and `view_frame()`
|
|
46
|
+
- a fixed row height and a table as high as a full page (default 10 rows) or
|
|
47
|
+
as all rows of a smaller diff, so filtering never makes it jump
|
|
48
|
+
- exact display of integers beyond 2^53 and of decimals a float cannot hold
|
|
49
|
+
- Demo notebook that tours every feature.
|
|
50
|
+
- Tests: unit tests, property-based tests (Hypothesis) and browser tests of the
|
|
51
|
+
widget (Playwright).
|
|
52
|
+
- Tooling: ruff and ty for Python, Biome and a strict JSDoc type check with
|
|
53
|
+
TypeScript for the frontend; all run in CI.
|
|
54
|
+
- Release workflow, started by hand with the version: runs the full CI, builds
|
|
55
|
+
and checks the distributions, publishes them to PyPI via trusted publishing
|
|
56
|
+
(with attestations), then tags the commit and creates the GitHub release. The
|
|
57
|
+
version comes from the git tag (hatch-vcs), so nothing needs to be bumped.
|
|
58
|
+
- ISC license.
|
|
59
|
+
|
|
60
|
+
[Unreleased]: https://github.com/oberbichler/dference/compare/v0.1.0...HEAD
|
|
61
|
+
[0.1.0]: https://github.com/oberbichler/dference/releases/tag/v0.1.0
|
dference-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
ISC License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Thomas Oberbichler
|
|
4
|
+
|
|
5
|
+
Permission to use, copy, modify, and/or distribute this software for any
|
|
6
|
+
purpose with or without fee is hereby granted, provided that the above
|
|
7
|
+
copyright notice and this permission notice appear in all copies.
|
|
8
|
+
|
|
9
|
+
THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES
|
|
10
|
+
WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF
|
|
11
|
+
MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR
|
|
12
|
+
ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES
|
|
13
|
+
WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN
|
|
14
|
+
ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF
|
|
15
|
+
OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE.
|
dference-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,283 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: dference
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Compare two DataFrames by key and explore the differences in an interactive widget for marimo and Jupyter.
|
|
5
|
+
Project-URL: Homepage, https://github.com/oberbichler/dference
|
|
6
|
+
Project-URL: Repository, https://github.com/oberbichler/dference
|
|
7
|
+
Project-URL: Issues, https://github.com/oberbichler/dference/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/oberbichler/dference/blob/main/CHANGELOG.md
|
|
9
|
+
Author-email: Thomas Oberbichler <thomas.oberbichler@gmail.com>
|
|
10
|
+
License-Expression: ISC
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: anywidget,compare,data-quality,dataframe,diff,jupyter,marimo,pandas,polars
|
|
13
|
+
Classifier: Development Status :: 4 - Beta
|
|
14
|
+
Classifier: Framework :: Jupyter
|
|
15
|
+
Classifier: Intended Audience :: Developers
|
|
16
|
+
Classifier: Intended Audience :: Science/Research
|
|
17
|
+
Classifier: Operating System :: OS Independent
|
|
18
|
+
Classifier: Programming Language :: Python :: 3
|
|
19
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
23
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
24
|
+
Classifier: Topic :: Scientific/Engineering :: Information Analysis
|
|
25
|
+
Classifier: Typing :: Typed
|
|
26
|
+
Requires-Python: >=3.11
|
|
27
|
+
Requires-Dist: anywidget>=0.9.13
|
|
28
|
+
Requires-Dist: polars>=1.30
|
|
29
|
+
Requires-Dist: traitlets>=5.9
|
|
30
|
+
Provides-Extra: pandas
|
|
31
|
+
Requires-Dist: pandas>=2.1; extra == 'pandas'
|
|
32
|
+
Requires-Dist: pyarrow>=15; extra == 'pandas'
|
|
33
|
+
Description-Content-Type: text/markdown
|
|
34
|
+
|
|
35
|
+
<!-- absolute URL: relative image paths do not render on PyPI -->
|
|
36
|
+
<p align="center">
|
|
37
|
+
<img src="https://raw.githubusercontent.com/oberbichler/dference/main/docs/logo.svg" alt="dference" width="480">
|
|
38
|
+
</p>
|
|
39
|
+
|
|
40
|
+
<p align="center">
|
|
41
|
+
<a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/v/dference" alt="PyPI"></a>
|
|
42
|
+
<a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/pyversions/dference" alt="Python versions"></a>
|
|
43
|
+
<a href="https://pypistats.org/packages/dference"><img src="https://img.shields.io/pypi/dm/dference" alt="Downloads"></a>
|
|
44
|
+
<a href="https://github.com/oberbichler/dference/actions/workflows/ci.yml"><img src="https://github.com/oberbichler/dference/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
|
|
45
|
+
<a href="https://github.com/oberbichler/dference/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-ISC-blue" alt="License: ISC"></a>
|
|
46
|
+
<br>
|
|
47
|
+
<a href="https://github.com/astral-sh/uv"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json" alt="uv"></a>
|
|
48
|
+
<a href="https://github.com/astral-sh/ruff"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json" alt="Ruff"></a>
|
|
49
|
+
<a href="https://github.com/astral-sh/ty"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ty/main/assets/badge/v0.json" alt="ty"></a>
|
|
50
|
+
<a href="https://molab.marimo.io/github/oberbichler/dference/blob/main/examples/demo.py"><img src="https://marimo.io/molab-shield.svg" alt="Open the demo in molab"></a>
|
|
51
|
+
</p>
|
|
52
|
+
|
|
53
|
+
**Compare two DataFrames by key and explore every difference – in an interactive
|
|
54
|
+
widget for [marimo](https://marimo.io) and Jupyter, or as plain polars tables.**
|
|
55
|
+
|
|
56
|
+
Rows are matched on a key (one or several columns) and classified as
|
|
57
|
+
|
|
58
|
+
| Status | Meaning |
|
|
59
|
+
| --- | --- |
|
|
60
|
+
| `equal` | key on both sides, all compared values equal |
|
|
61
|
+
| `mismatch` | key on both sides, at least one value differs |
|
|
62
|
+
| `missing_right` | key only in the left frame (shown as `Only in {left_name}`) |
|
|
63
|
+
| `missing_left` | key only in the right frame (shown as `Only in {right_name}`) |
|
|
64
|
+
|
|
65
|
+
The widget is modelled on `marimo.ui.table`: paging, sorting, search, a status
|
|
66
|
+
filter, marimo-style column filters, per-column match rates, a side-by-side
|
|
67
|
+
detail view, invisible-character highlighting and CSV export. It stays fast on
|
|
68
|
+
millions of rows because filtering, sorting and paging run in polars on the
|
|
69
|
+
Python side – the browser only ever receives the visible page.
|
|
70
|
+
|
|
71
|
+
## Installation
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
uv add dference # or: pip install dference
|
|
75
|
+
uv add "dference[pandas]" # if your inputs are pandas DataFrames
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
`dference` depends on `polars`, `anywidget` and `traitlets`. Inputs can be polars
|
|
79
|
+
(eager or lazy), pandas, pyarrow, or anything implementing the Arrow PyCapsule
|
|
80
|
+
interface or `to_polars()` (DuckDB relations, narwhals, …).
|
|
81
|
+
|
|
82
|
+
## Quick start
|
|
83
|
+
|
|
84
|
+
### marimo
|
|
85
|
+
|
|
86
|
+
```python
|
|
87
|
+
import marimo as mo
|
|
88
|
+
from dference import DataFrameDiff
|
|
89
|
+
|
|
90
|
+
diff = DataFrameDiff(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
|
|
91
|
+
view = mo.ui.anywidget(diff)
|
|
92
|
+
view
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
In another cell, work with the rows checked in the widget. The cell re-runs
|
|
96
|
+
when the selection changes – and only then; paging, sorting and filtering
|
|
97
|
+
never trigger re-execution:
|
|
98
|
+
|
|
99
|
+
```python
|
|
100
|
+
view.value["selected_ids"] # reactive dependency
|
|
101
|
+
diff.selected_frame() # the checked rows as a polars DataFrame
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
### Jupyter
|
|
105
|
+
|
|
106
|
+
```python
|
|
107
|
+
from dference import DataFrameDiff
|
|
108
|
+
|
|
109
|
+
DataFrameDiff(crm, erp, key=["region", "customer_id"])
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
### Without a UI
|
|
113
|
+
|
|
114
|
+
```python
|
|
115
|
+
import dference
|
|
116
|
+
|
|
117
|
+
result = dference.compare(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
|
|
118
|
+
|
|
119
|
+
result.summary # Summary(equal=158, mismatch=78, missing_left=10, missing_right=8, ...)
|
|
120
|
+
result.frame() # wide: key, status, differing, "<col> [CRM]", "<col> [ERP]", ...
|
|
121
|
+
result.frame("mismatch")
|
|
122
|
+
result.mismatches() # long: one row per (key, column) that differs
|
|
123
|
+
result.column_stats() # per column: dtypes, mismatches, equal/mismatch share
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
## How values are compared
|
|
127
|
+
|
|
128
|
+
- **Exact equality.** There is deliberately no numeric tolerance. If small float
|
|
129
|
+
differences should not count, round both frames first, e.g.
|
|
130
|
+
`df.with_columns(cs.float().round(2))`.
|
|
131
|
+
- **Missing values are equal to each other**: `null == null`, and in float
|
|
132
|
+
columns `NaN` counts as missing too (pandas uses NaN as its null marker, so
|
|
133
|
+
a frame from pandas and one from Arrow would otherwise disagree).
|
|
134
|
+
- **Same column, same dtype.** By default (`strict=True`) a column must have
|
|
135
|
+
the identical dtype on both sides.
|
|
136
|
+
- **`strict=False`** aligns differences that lose nothing – the same values
|
|
137
|
+
stored differently: integers of different width, `Float32` vs. `Float64`,
|
|
138
|
+
decimals of different precision, `Datetime` / `Duration` in different time
|
|
139
|
+
units (same time zone), text vs. `Categorical` / `Enum`, and an all-null
|
|
140
|
+
column vs. anything. `ColumnInfo.compare_dtype` tells you which dtype was
|
|
141
|
+
used.
|
|
142
|
+
- **Everything else is always your decision.** Int vs. float, `Date` vs. `Datetime`,
|
|
143
|
+
naive vs. time-zone aware, different time zones, text vs. numbers, different
|
|
144
|
+
nested types: `compare()` raises one `ValueError` that lists every such column
|
|
145
|
+
(keys included) with a cast to fix it, e.g.
|
|
146
|
+
`right = right.with_columns(pl.col("qty").cast(pl.Int64))`. Note that pandas
|
|
147
|
+
has no date dtype – a `Date` column from pandas arrives as a `Datetime`.
|
|
148
|
+
- **Keys** must be unique on each side; null keys match null keys.
|
|
149
|
+
- **Columns present on one side only** are shown but not compared. Use
|
|
150
|
+
`ignore_columns=[...]` to leave columns out entirely.
|
|
151
|
+
|
|
152
|
+
## Widget features
|
|
153
|
+
|
|
154
|
+
- Status overview on a shared scale (equal / mismatch / only left / only
|
|
155
|
+
right, plus found vs. not found); filter by status from the menu in the
|
|
156
|
+
status column header.
|
|
157
|
+
- Each side has one colour and one chip – the first letter of its name
|
|
158
|
+
(`left_short=` / `right_short=` to override) – used everywhere: differences
|
|
159
|
+
show both values with equal weight, each marked by its side's chip, and rows
|
|
160
|
+
on one side only carry the chip of the side they exist on.
|
|
161
|
+
- Per-column match rate in the header; click `≠` or `=` to show only rows
|
|
162
|
+
that differ in, or are equal in, that column.
|
|
163
|
+
- Column menu per header: sort and type-aware filters (text, number range,
|
|
164
|
+
date range, boolean, null checks). For compared columns **a row matches if
|
|
165
|
+
the left or the right value matches**; a side that does not exist in a row
|
|
166
|
+
never matches.
|
|
167
|
+
- Invisible characters are made visible: `␣` leading/trailing/repeated spaces,
|
|
168
|
+
`→` tab, `↵` line feed, `⍽` no-break spaces, `ZWSP`, `BOM`, … Cells that
|
|
169
|
+
differ *only* in invisible characters are flagged.
|
|
170
|
+
- Side-by-side detail view with numeric deltas; step through the filtered rows
|
|
171
|
+
across page boundaries.
|
|
172
|
+
- CSV export of all rows matching the current filters.
|
|
173
|
+
|
|
174
|
+
## Python API
|
|
175
|
+
|
|
176
|
+
| Object | Purpose |
|
|
177
|
+
| --- | --- |
|
|
178
|
+
| `compare(left, right, key, *, left_name, right_name, ignore_columns, strict=True)` | Run a comparison, returns `DiffResult`. |
|
|
179
|
+
| `DiffResult.summary` | `Summary` with counts per status plus `found`, `not_found`, `total`. |
|
|
180
|
+
| `DiffResult.columns` | Tuple of `ColumnInfo` (kind, dtypes, mismatches, compare dtype). |
|
|
181
|
+
| `DiffResult.frame(status=None, *, rows=None)` | Wide result, optionally filtered by status or row ids. |
|
|
182
|
+
| `DiffResult.mismatches()` | Long format of all differing cells (values as text). |
|
|
183
|
+
| `DiffResult.column_stats()` | Match rates per compared column. |
|
|
184
|
+
| `DataFrameDiff(left, right, key, ...)` | The widget; same arguments as `compare` plus `left_short`, `right_short`, `page_size` (default 10). |
|
|
185
|
+
| `DataFrameDiff.from_result(result, ...)` | Widget for an existing `DiffResult`. |
|
|
186
|
+
| `DataFrameDiff.result` | The underlying `DiffResult`. |
|
|
187
|
+
| `DataFrameDiff.selected_ids` | Synced trait with the checked row ids. |
|
|
188
|
+
| `DataFrameDiff.selected_frame()` | Checked rows as a wide frame. |
|
|
189
|
+
| `DataFrameDiff.view_frame()` | Rows matching the widget's current filters (not reactive). |
|
|
190
|
+
| `Status` | `StrEnum` of the four statuses. |
|
|
191
|
+
|
|
192
|
+
## Performance
|
|
193
|
+
|
|
194
|
+
The comparison is one hash join plus vectorised expressions; the widget keeps
|
|
195
|
+
the ordered row ids of the last few views in a small cache, so paging through
|
|
196
|
+
a view is a slice. Timings from `benchmarks/bench.py` on a **single CPU core**
|
|
197
|
+
(polars parallelises, so more cores are faster):
|
|
198
|
+
|
|
199
|
+
| 5 million rows per side | time |
|
|
200
|
+
| --- | ---: |
|
|
201
|
+
| `compare` (join, difference flags, statistics) | 2.5 s |
|
|
202
|
+
| first page / next page | < 1 ms |
|
|
203
|
+
| filter by status | 19 ms |
|
|
204
|
+
| numeric range filter (left or right) | 45 ms |
|
|
205
|
+
| sort by a column | 0.4 s |
|
|
206
|
+
| full-text search across all columns | 0.4 s |
|
|
207
|
+
|
|
208
|
+
Memory: the joined frame holds both sides once, plus one boolean flag per
|
|
209
|
+
compared column.
|
|
210
|
+
|
|
211
|
+
## Limitations
|
|
212
|
+
|
|
213
|
+
- The widget needs a running Python kernel; in a static HTML export it shows no
|
|
214
|
+
rows.
|
|
215
|
+
- CSV export sends all matching rows to the browser in one message – filter
|
|
216
|
+
first for very large exports, or use `view_frame().write_csv(...)`.
|
|
217
|
+
- `view_frame()` reflects the filters last sent by the widget but is not a
|
|
218
|
+
reactive value in marimo.
|
|
219
|
+
|
|
220
|
+
## Development
|
|
221
|
+
|
|
222
|
+
```bash
|
|
223
|
+
git clone https://github.com/oberbichler/dference && cd dference
|
|
224
|
+
uv sync # creates .venv with all dev tools
|
|
225
|
+
|
|
226
|
+
uv run ruff format . # format
|
|
227
|
+
uv run ruff check . # lint
|
|
228
|
+
uv run ty check # type check
|
|
229
|
+
uv run pytest --cov # tests with coverage
|
|
230
|
+
uv run playwright install chromium # once: browser for the widget tests
|
|
231
|
+
uv run pytest tests/e2e # widget in a real browser (skipped without Chromium)
|
|
232
|
+
|
|
233
|
+
uv run marimo edit examples/demo.py # interactive demo
|
|
234
|
+
uv run python benchmarks/bench.py 1_000_000
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
The frontend is a single dependency-free ES module
|
|
238
|
+
(`src/dference/static/widget.js`) with its stylesheet next to it – no build step.
|
|
239
|
+
anywidget reloads both on change while you develop.
|
|
240
|
+
|
|
241
|
+
Node.js is only needed for frontend tooling, never at runtime:
|
|
242
|
+
|
|
243
|
+
```bash
|
|
244
|
+
npm ci # Biome + TypeScript (dev only)
|
|
245
|
+
|
|
246
|
+
npm run fix # format + auto-fix JS, CSS and JSON (Biome)
|
|
247
|
+
npm run check # lint (Biome) and type-check the JSDoc (tsc)
|
|
248
|
+
```
|
|
249
|
+
|
|
250
|
+
Types live in JSDoc comments and are checked with `tsc --checkJs` in strict
|
|
251
|
+
mode (see `tsconfig.json`); the widget API is typed via `@anywidget/types`.
|
|
252
|
+
Nothing is compiled – the file that ships is the file you edit.
|
|
253
|
+
|
|
254
|
+
### Releasing
|
|
255
|
+
|
|
256
|
+
The version is derived from the git tag (`v1.2.3`) by
|
|
257
|
+
[hatch-vcs](https://github.com/ofek/hatch-vcs) – there is no version to bump in
|
|
258
|
+
`pyproject.toml` or `uv.lock`. A checkout between releases reports a dev version
|
|
259
|
+
such as `0.1.1.dev3+g94570d5`.
|
|
260
|
+
|
|
261
|
+
For every release:
|
|
262
|
+
|
|
263
|
+
1. Rename `## [Unreleased]` in `CHANGELOG.md` to `## [x.y.z] - <date>` (add a new
|
|
264
|
+
empty `Unreleased` section above it) and merge that to `main`.
|
|
265
|
+
2. Run *Actions → Release → Run workflow* on `main` with the version `x.y.z`, or
|
|
266
|
+
`gh workflow run release.yml -f version=x.y.z`.
|
|
267
|
+
|
|
268
|
+
The workflow refuses versions that already exist as tag, release or on PyPI,
|
|
269
|
+
runs the full CI on the commit, builds with that version, publishes to PyPI via
|
|
270
|
+
[trusted publishing](https://docs.pypi.org/trusted-publishers/), and only then
|
|
271
|
+
tags the commit `vx.y.z` and creates the GitHub release with the changelog
|
|
272
|
+
section as notes and the wheel and sdist attached. If a step fails, *Re-run
|
|
273
|
+
failed jobs* continues from there.
|
|
274
|
+
|
|
275
|
+
Once, before the first release: on PyPI add a
|
|
276
|
+
[pending trusted publisher](https://docs.pypi.org/trusted-publishers/creating-a-project-through-oidc/)
|
|
277
|
+
for the project `dference` – owner `oberbichler`, repository `dference`,
|
|
278
|
+
workflow `release.yml`, environment `pypi` – and create the environment `pypi`
|
|
279
|
+
in the GitHub repository settings (optionally with required reviewers).
|
|
280
|
+
|
|
281
|
+
## License
|
|
282
|
+
|
|
283
|
+
ISC © Thomas Oberbichler
|
dference-0.1.0/README.md
ADDED
|
@@ -0,0 +1,249 @@
|
|
|
1
|
+
<!-- absolute URL: relative image paths do not render on PyPI -->
|
|
2
|
+
<p align="center">
|
|
3
|
+
<img src="https://raw.githubusercontent.com/oberbichler/dference/main/docs/logo.svg" alt="dference" width="480">
|
|
4
|
+
</p>
|
|
5
|
+
|
|
6
|
+
<p align="center">
|
|
7
|
+
<a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/v/dference" alt="PyPI"></a>
|
|
8
|
+
<a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/pyversions/dference" alt="Python versions"></a>
|
|
9
|
+
<a href="https://pypistats.org/packages/dference"><img src="https://img.shields.io/pypi/dm/dference" alt="Downloads"></a>
|
|
10
|
+
<a href="https://github.com/oberbichler/dference/actions/workflows/ci.yml"><img src="https://github.com/oberbichler/dference/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
|
|
11
|
+
<a href="https://github.com/oberbichler/dference/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-ISC-blue" alt="License: ISC"></a>
|
|
12
|
+
<br>
|
|
13
|
+
<a href="https://github.com/astral-sh/uv"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json" alt="uv"></a>
|
|
14
|
+
<a href="https://github.com/astral-sh/ruff"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json" alt="Ruff"></a>
|
|
15
|
+
<a href="https://github.com/astral-sh/ty"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ty/main/assets/badge/v0.json" alt="ty"></a>
|
|
16
|
+
<a href="https://molab.marimo.io/github/oberbichler/dference/blob/main/examples/demo.py"><img src="https://marimo.io/molab-shield.svg" alt="Open the demo in molab"></a>
|
|
17
|
+
</p>
|
|
18
|
+
|
|
19
|
+
**Compare two DataFrames by key and explore every difference – in an interactive
|
|
20
|
+
widget for [marimo](https://marimo.io) and Jupyter, or as plain polars tables.**
|
|
21
|
+
|
|
22
|
+
Rows are matched on a key (one or several columns) and classified as
|
|
23
|
+
|
|
24
|
+
| Status | Meaning |
|
|
25
|
+
| --- | --- |
|
|
26
|
+
| `equal` | key on both sides, all compared values equal |
|
|
27
|
+
| `mismatch` | key on both sides, at least one value differs |
|
|
28
|
+
| `missing_right` | key only in the left frame (shown as `Only in {left_name}`) |
|
|
29
|
+
| `missing_left` | key only in the right frame (shown as `Only in {right_name}`) |
|
|
30
|
+
|
|
31
|
+
The widget is modelled on `marimo.ui.table`: paging, sorting, search, a status
|
|
32
|
+
filter, marimo-style column filters, per-column match rates, a side-by-side
|
|
33
|
+
detail view, invisible-character highlighting and CSV export. It stays fast on
|
|
34
|
+
millions of rows because filtering, sorting and paging run in polars on the
|
|
35
|
+
Python side – the browser only ever receives the visible page.
|
|
36
|
+
|
|
37
|
+
## Installation
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
uv add dference # or: pip install dference
|
|
41
|
+
uv add "dference[pandas]" # if your inputs are pandas DataFrames
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
`dference` depends on `polars`, `anywidget` and `traitlets`. Inputs can be polars
|
|
45
|
+
(eager or lazy), pandas, pyarrow, or anything implementing the Arrow PyCapsule
|
|
46
|
+
interface or `to_polars()` (DuckDB relations, narwhals, …).
|
|
47
|
+
|
|
48
|
+
## Quick start
|
|
49
|
+
|
|
50
|
+
### marimo
|
|
51
|
+
|
|
52
|
+
```python
|
|
53
|
+
import marimo as mo
|
|
54
|
+
from dference import DataFrameDiff
|
|
55
|
+
|
|
56
|
+
diff = DataFrameDiff(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
|
|
57
|
+
view = mo.ui.anywidget(diff)
|
|
58
|
+
view
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
In another cell, work with the rows checked in the widget. The cell re-runs
|
|
62
|
+
when the selection changes – and only then; paging, sorting and filtering
|
|
63
|
+
never trigger re-execution:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
view.value["selected_ids"] # reactive dependency
|
|
67
|
+
diff.selected_frame() # the checked rows as a polars DataFrame
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
### Jupyter
|
|
71
|
+
|
|
72
|
+
```python
|
|
73
|
+
from dference import DataFrameDiff
|
|
74
|
+
|
|
75
|
+
DataFrameDiff(crm, erp, key=["region", "customer_id"])
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
### Without a UI
|
|
79
|
+
|
|
80
|
+
```python
|
|
81
|
+
import dference
|
|
82
|
+
|
|
83
|
+
result = dference.compare(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
|
|
84
|
+
|
|
85
|
+
result.summary # Summary(equal=158, mismatch=78, missing_left=10, missing_right=8, ...)
|
|
86
|
+
result.frame() # wide: key, status, differing, "<col> [CRM]", "<col> [ERP]", ...
|
|
87
|
+
result.frame("mismatch")
|
|
88
|
+
result.mismatches() # long: one row per (key, column) that differs
|
|
89
|
+
result.column_stats() # per column: dtypes, mismatches, equal/mismatch share
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
## How values are compared
|
|
93
|
+
|
|
94
|
+
- **Exact equality.** There is deliberately no numeric tolerance. If small float
|
|
95
|
+
differences should not count, round both frames first, e.g.
|
|
96
|
+
`df.with_columns(cs.float().round(2))`.
|
|
97
|
+
- **Missing values are equal to each other**: `null == null`, and in float
|
|
98
|
+
columns `NaN` counts as missing too (pandas uses NaN as its null marker, so
|
|
99
|
+
a frame from pandas and one from Arrow would otherwise disagree).
|
|
100
|
+
- **Same column, same dtype.** By default (`strict=True`) a column must have
|
|
101
|
+
the identical dtype on both sides.
|
|
102
|
+
- **`strict=False`** aligns differences that lose nothing – the same values
|
|
103
|
+
stored differently: integers of different width, `Float32` vs. `Float64`,
|
|
104
|
+
decimals of different precision, `Datetime` / `Duration` in different time
|
|
105
|
+
units (same time zone), text vs. `Categorical` / `Enum`, and an all-null
|
|
106
|
+
column vs. anything. `ColumnInfo.compare_dtype` tells you which dtype was
|
|
107
|
+
used.
|
|
108
|
+
- **Everything else is always your decision.** Int vs. float, `Date` vs. `Datetime`,
|
|
109
|
+
naive vs. time-zone aware, different time zones, text vs. numbers, different
|
|
110
|
+
nested types: `compare()` raises one `ValueError` that lists every such column
|
|
111
|
+
(keys included) with a cast to fix it, e.g.
|
|
112
|
+
`right = right.with_columns(pl.col("qty").cast(pl.Int64))`. Note that pandas
|
|
113
|
+
has no date dtype – a `Date` column from pandas arrives as a `Datetime`.
|
|
114
|
+
- **Keys** must be unique on each side; null keys match null keys.
|
|
115
|
+
- **Columns present on one side only** are shown but not compared. Use
|
|
116
|
+
`ignore_columns=[...]` to leave columns out entirely.
|
|
117
|
+
|
|
118
|
+
## Widget features
|
|
119
|
+
|
|
120
|
+
- Status overview on a shared scale (equal / mismatch / only left / only
|
|
121
|
+
right, plus found vs. not found); filter by status from the menu in the
|
|
122
|
+
status column header.
|
|
123
|
+
- Each side has one colour and one chip – the first letter of its name
|
|
124
|
+
(`left_short=` / `right_short=` to override) – used everywhere: differences
|
|
125
|
+
show both values with equal weight, each marked by its side's chip, and rows
|
|
126
|
+
on one side only carry the chip of the side they exist on.
|
|
127
|
+
- Per-column match rate in the header; click `≠` or `=` to show only rows
|
|
128
|
+
that differ in, or are equal in, that column.
|
|
129
|
+
- Column menu per header: sort and type-aware filters (text, number range,
|
|
130
|
+
date range, boolean, null checks). For compared columns **a row matches if
|
|
131
|
+
the left or the right value matches**; a side that does not exist in a row
|
|
132
|
+
never matches.
|
|
133
|
+
- Invisible characters are made visible: `␣` leading/trailing/repeated spaces,
|
|
134
|
+
`→` tab, `↵` line feed, `⍽` no-break spaces, `ZWSP`, `BOM`, … Cells that
|
|
135
|
+
differ *only* in invisible characters are flagged.
|
|
136
|
+
- Side-by-side detail view with numeric deltas; step through the filtered rows
|
|
137
|
+
across page boundaries.
|
|
138
|
+
- CSV export of all rows matching the current filters.
|
|
139
|
+
|
|
140
|
+
## Python API
|
|
141
|
+
|
|
142
|
+
| Object | Purpose |
|
|
143
|
+
| --- | --- |
|
|
144
|
+
| `compare(left, right, key, *, left_name, right_name, ignore_columns, strict=True)` | Run a comparison, returns `DiffResult`. |
|
|
145
|
+
| `DiffResult.summary` | `Summary` with counts per status plus `found`, `not_found`, `total`. |
|
|
146
|
+
| `DiffResult.columns` | Tuple of `ColumnInfo` (kind, dtypes, mismatches, compare dtype). |
|
|
147
|
+
| `DiffResult.frame(status=None, *, rows=None)` | Wide result, optionally filtered by status or row ids. |
|
|
148
|
+
| `DiffResult.mismatches()` | Long format of all differing cells (values as text). |
|
|
149
|
+
| `DiffResult.column_stats()` | Match rates per compared column. |
|
|
150
|
+
| `DataFrameDiff(left, right, key, ...)` | The widget; same arguments as `compare` plus `left_short`, `right_short`, `page_size` (default 10). |
|
|
151
|
+
| `DataFrameDiff.from_result(result, ...)` | Widget for an existing `DiffResult`. |
|
|
152
|
+
| `DataFrameDiff.result` | The underlying `DiffResult`. |
|
|
153
|
+
| `DataFrameDiff.selected_ids` | Synced trait with the checked row ids. |
|
|
154
|
+
| `DataFrameDiff.selected_frame()` | Checked rows as a wide frame. |
|
|
155
|
+
| `DataFrameDiff.view_frame()` | Rows matching the widget's current filters (not reactive). |
|
|
156
|
+
| `Status` | `StrEnum` of the four statuses. |
|
|
157
|
+
|
|
158
|
+
## Performance
|
|
159
|
+
|
|
160
|
+
The comparison is one hash join plus vectorised expressions; the widget keeps
|
|
161
|
+
the ordered row ids of the last few views in a small cache, so paging through
|
|
162
|
+
a view is a slice. Timings from `benchmarks/bench.py` on a **single CPU core**
|
|
163
|
+
(polars parallelises, so more cores are faster):
|
|
164
|
+
|
|
165
|
+
| 5 million rows per side | time |
|
|
166
|
+
| --- | ---: |
|
|
167
|
+
| `compare` (join, difference flags, statistics) | 2.5 s |
|
|
168
|
+
| first page / next page | < 1 ms |
|
|
169
|
+
| filter by status | 19 ms |
|
|
170
|
+
| numeric range filter (left or right) | 45 ms |
|
|
171
|
+
| sort by a column | 0.4 s |
|
|
172
|
+
| full-text search across all columns | 0.4 s |
|
|
173
|
+
|
|
174
|
+
Memory: the joined frame holds both sides once, plus one boolean flag per
|
|
175
|
+
compared column.
|
|
176
|
+
|
|
177
|
+
## Limitations
|
|
178
|
+
|
|
179
|
+
- The widget needs a running Python kernel; in a static HTML export it shows no
|
|
180
|
+
rows.
|
|
181
|
+
- CSV export sends all matching rows to the browser in one message – filter
|
|
182
|
+
first for very large exports, or use `view_frame().write_csv(...)`.
|
|
183
|
+
- `view_frame()` reflects the filters last sent by the widget but is not a
|
|
184
|
+
reactive value in marimo.
|
|
185
|
+
|
|
186
|
+
## Development
|
|
187
|
+
|
|
188
|
+
```bash
|
|
189
|
+
git clone https://github.com/oberbichler/dference && cd dference
|
|
190
|
+
uv sync # creates .venv with all dev tools
|
|
191
|
+
|
|
192
|
+
uv run ruff format . # format
|
|
193
|
+
uv run ruff check . # lint
|
|
194
|
+
uv run ty check # type check
|
|
195
|
+
uv run pytest --cov # tests with coverage
|
|
196
|
+
uv run playwright install chromium # once: browser for the widget tests
|
|
197
|
+
uv run pytest tests/e2e # widget in a real browser (skipped without Chromium)
|
|
198
|
+
|
|
199
|
+
uv run marimo edit examples/demo.py # interactive demo
|
|
200
|
+
uv run python benchmarks/bench.py 1_000_000
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
The frontend is a single dependency-free ES module
|
|
204
|
+
(`src/dference/static/widget.js`) with its stylesheet next to it – no build step.
|
|
205
|
+
anywidget reloads both on change while you develop.
|
|
206
|
+
|
|
207
|
+
Node.js is only needed for frontend tooling, never at runtime:
|
|
208
|
+
|
|
209
|
+
```bash
|
|
210
|
+
npm ci # Biome + TypeScript (dev only)
|
|
211
|
+
|
|
212
|
+
npm run fix # format + auto-fix JS, CSS and JSON (Biome)
|
|
213
|
+
npm run check # lint (Biome) and type-check the JSDoc (tsc)
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
Types live in JSDoc comments and are checked with `tsc --checkJs` in strict
|
|
217
|
+
mode (see `tsconfig.json`); the widget API is typed via `@anywidget/types`.
|
|
218
|
+
Nothing is compiled – the file that ships is the file you edit.
|
|
219
|
+
|
|
220
|
+
### Releasing
|
|
221
|
+
|
|
222
|
+
The version is derived from the git tag (`v1.2.3`) by
|
|
223
|
+
[hatch-vcs](https://github.com/ofek/hatch-vcs) – there is no version to bump in
|
|
224
|
+
`pyproject.toml` or `uv.lock`. A checkout between releases reports a dev version
|
|
225
|
+
such as `0.1.1.dev3+g94570d5`.
|
|
226
|
+
|
|
227
|
+
For every release:
|
|
228
|
+
|
|
229
|
+
1. Rename `## [Unreleased]` in `CHANGELOG.md` to `## [x.y.z] - <date>` (add a new
|
|
230
|
+
empty `Unreleased` section above it) and merge that to `main`.
|
|
231
|
+
2. Run *Actions → Release → Run workflow* on `main` with the version `x.y.z`, or
|
|
232
|
+
`gh workflow run release.yml -f version=x.y.z`.
|
|
233
|
+
|
|
234
|
+
The workflow refuses versions that already exist as tag, release or on PyPI,
|
|
235
|
+
runs the full CI on the commit, builds with that version, publishes to PyPI via
|
|
236
|
+
[trusted publishing](https://docs.pypi.org/trusted-publishers/), and only then
|
|
237
|
+
tags the commit `vx.y.z` and creates the GitHub release with the changelog
|
|
238
|
+
section as notes and the wheel and sdist attached. If a step fails, *Re-run
|
|
239
|
+
failed jobs* continues from there.
|
|
240
|
+
|
|
241
|
+
Once, before the first release: on PyPI add a
|
|
242
|
+
[pending trusted publisher](https://docs.pypi.org/trusted-publishers/creating-a-project-through-oidc/)
|
|
243
|
+
for the project `dference` – owner `oberbichler`, repository `dference`,
|
|
244
|
+
workflow `release.yml`, environment `pypi` – and create the environment `pypi`
|
|
245
|
+
in the GitHub repository settings (optionally with required reviewers).
|
|
246
|
+
|
|
247
|
+
## License
|
|
248
|
+
|
|
249
|
+
ISC © Thomas Oberbichler
|