dference 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,26 @@
1
+ # Python
2
+ __pycache__/
3
+ *.py[cod]
4
+ .venv/
5
+ dist/
6
+ build/
7
+ *.egg-info/
8
+
9
+ # Frontend tooling
10
+ node_modules/
11
+
12
+ # Tooling caches
13
+ .ruff_cache/
14
+ .pytest_cache/
15
+ .coverage
16
+ coverage.xml
17
+ htmlcov/
18
+
19
+ # marimo
20
+ __marimo__/
21
+ .marimo/
22
+
23
+ # OS / editors
24
+ .DS_Store
25
+ .idea/
26
+ .vscode/
@@ -0,0 +1,61 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the project uses
5
+ [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [Unreleased]
8
+
9
+ ## [0.1.0] - 2026-10-06
10
+
11
+ First release.
12
+
13
+ ### Added
14
+
15
+ - `compare()` – match two DataFrames on a (composite) key with a single polars
16
+ join and classify every row as equal, mismatch, only in left
17
+ (`missing_right`) or only in right (`missing_left`).
18
+ - `DiffResult` with `summary`, `columns`, `frame()`, `mismatches()` and
19
+ `column_stats()`.
20
+ - Exact comparison without tolerance; `null` and `NaN` count as equal missing
21
+ values.
22
+ - Strict dtypes: `strict=True` (the default) requires identical dtypes on both
23
+ sides; `strict=False` aligns lossless differences (integer width,
24
+ `Float32`/`Float64`, decimal precision, time units within one time zone, text
25
+ vs. `Categorical`/`Enum`, all-null columns). Anything else – including naive
26
+ vs. time-zone-aware and different time zones – raises one error listing each
27
+ column with a cast.
28
+ - Input support for polars (eager and lazy), pandas, pyarrow and objects with
29
+ the Arrow PyCapsule interface or `to_polars()`.
30
+ - `DataFrameDiff` anywidget for marimo and Jupyter. Filtering, sorting and
31
+ paging run in polars on the Python side; the browser only receives the
32
+ visible page. It has:
33
+ - a status overview on a shared scale, ordered left before right
34
+ - a status filter in the header of the status column; an active filter shows
35
+ as a removable pill in the toolbar
36
+ - quick filters `≠` and `=` per column (only rows that differ in, or are
37
+ equal in, that column) next to per-column match rates
38
+ - column menus with sorting and type-aware filters, where a row matches if
39
+ the left *or* the right value matches
40
+ - one colour and one chip (the side's short name) per side, used for
41
+ difference markers, statuses, one-sided columns, overview and detail view
42
+ - a side-by-side detail view with numeric deltas, stepping across pages
43
+ - invisible-character highlighting
44
+ - search, CSV export of the filtered rows, row selection synced to Python
45
+ (`selected_ids`, `selected_frame()`) and `view_frame()`
46
+ - a fixed row height and a table as high as a full page (default 10 rows) or
47
+ as all rows of a smaller diff, so filtering never makes it jump
48
+ - exact display of integers beyond 2^53 and of decimals a float cannot hold
49
+ - Demo notebook that tours every feature.
50
+ - Tests: unit tests, property-based tests (Hypothesis) and browser tests of the
51
+ widget (Playwright).
52
+ - Tooling: ruff and ty for Python, Biome and a strict JSDoc type check with
53
+ TypeScript for the frontend; all run in CI.
54
+ - Release workflow, started by hand with the version: runs the full CI, builds
55
+ and checks the distributions, publishes them to PyPI via trusted publishing
56
+ (with attestations), then tags the commit and creates the GitHub release. The
57
+ version comes from the git tag (hatch-vcs), so nothing needs to be bumped.
58
+ - ISC license.
59
+
60
+ [Unreleased]: https://github.com/oberbichler/dference/compare/v0.1.0...HEAD
61
+ [0.1.0]: https://github.com/oberbichler/dference/releases/tag/v0.1.0
dference-0.1.0/LICENSE ADDED
@@ -0,0 +1,15 @@
1
+ ISC License
2
+
3
+ Copyright (c) 2026 Thomas Oberbichler
4
+
5
+ Permission to use, copy, modify, and/or distribute this software for any
6
+ purpose with or without fee is hereby granted, provided that the above
7
+ copyright notice and this permission notice appear in all copies.
8
+
9
+ THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES
10
+ WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF
11
+ MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR
12
+ ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES
13
+ WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN
14
+ ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF
15
+ OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE.
@@ -0,0 +1,283 @@
1
+ Metadata-Version: 2.5
2
+ Name: dference
3
+ Version: 0.1.0
4
+ Summary: Compare two DataFrames by key and explore the differences in an interactive widget for marimo and Jupyter.
5
+ Project-URL: Homepage, https://github.com/oberbichler/dference
6
+ Project-URL: Repository, https://github.com/oberbichler/dference
7
+ Project-URL: Issues, https://github.com/oberbichler/dference/issues
8
+ Project-URL: Changelog, https://github.com/oberbichler/dference/blob/main/CHANGELOG.md
9
+ Author-email: Thomas Oberbichler <thomas.oberbichler@gmail.com>
10
+ License-Expression: ISC
11
+ License-File: LICENSE
12
+ Keywords: anywidget,compare,data-quality,dataframe,diff,jupyter,marimo,pandas,polars
13
+ Classifier: Development Status :: 4 - Beta
14
+ Classifier: Framework :: Jupyter
15
+ Classifier: Intended Audience :: Developers
16
+ Classifier: Intended Audience :: Science/Research
17
+ Classifier: Operating System :: OS Independent
18
+ Classifier: Programming Language :: Python :: 3
19
+ Classifier: Programming Language :: Python :: 3 :: Only
20
+ Classifier: Programming Language :: Python :: 3.11
21
+ Classifier: Programming Language :: Python :: 3.12
22
+ Classifier: Programming Language :: Python :: 3.13
23
+ Classifier: Programming Language :: Python :: 3.14
24
+ Classifier: Topic :: Scientific/Engineering :: Information Analysis
25
+ Classifier: Typing :: Typed
26
+ Requires-Python: >=3.11
27
+ Requires-Dist: anywidget>=0.9.13
28
+ Requires-Dist: polars>=1.30
29
+ Requires-Dist: traitlets>=5.9
30
+ Provides-Extra: pandas
31
+ Requires-Dist: pandas>=2.1; extra == 'pandas'
32
+ Requires-Dist: pyarrow>=15; extra == 'pandas'
33
+ Description-Content-Type: text/markdown
34
+
35
+ <!-- absolute URL: relative image paths do not render on PyPI -->
36
+ <p align="center">
37
+ <img src="https://raw.githubusercontent.com/oberbichler/dference/main/docs/logo.svg" alt="dference" width="480">
38
+ </p>
39
+
40
+ <p align="center">
41
+ <a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/v/dference" alt="PyPI"></a>
42
+ <a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/pyversions/dference" alt="Python versions"></a>
43
+ <a href="https://pypistats.org/packages/dference"><img src="https://img.shields.io/pypi/dm/dference" alt="Downloads"></a>
44
+ <a href="https://github.com/oberbichler/dference/actions/workflows/ci.yml"><img src="https://github.com/oberbichler/dference/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
45
+ <a href="https://github.com/oberbichler/dference/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-ISC-blue" alt="License: ISC"></a>
46
+ <br>
47
+ <a href="https://github.com/astral-sh/uv"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json" alt="uv"></a>
48
+ <a href="https://github.com/astral-sh/ruff"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json" alt="Ruff"></a>
49
+ <a href="https://github.com/astral-sh/ty"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ty/main/assets/badge/v0.json" alt="ty"></a>
50
+ <a href="https://molab.marimo.io/github/oberbichler/dference/blob/main/examples/demo.py"><img src="https://marimo.io/molab-shield.svg" alt="Open the demo in molab"></a>
51
+ </p>
52
+
53
+ **Compare two DataFrames by key and explore every difference – in an interactive
54
+ widget for [marimo](https://marimo.io) and Jupyter, or as plain polars tables.**
55
+
56
+ Rows are matched on a key (one or several columns) and classified as
57
+
58
+ | Status | Meaning |
59
+ | --- | --- |
60
+ | `equal` | key on both sides, all compared values equal |
61
+ | `mismatch` | key on both sides, at least one value differs |
62
+ | `missing_right` | key only in the left frame (shown as `Only in {left_name}`) |
63
+ | `missing_left` | key only in the right frame (shown as `Only in {right_name}`) |
64
+
65
+ The widget is modelled on `marimo.ui.table`: paging, sorting, search, a status
66
+ filter, marimo-style column filters, per-column match rates, a side-by-side
67
+ detail view, invisible-character highlighting and CSV export. It stays fast on
68
+ millions of rows because filtering, sorting and paging run in polars on the
69
+ Python side – the browser only ever receives the visible page.
70
+
71
+ ## Installation
72
+
73
+ ```bash
74
+ uv add dference # or: pip install dference
75
+ uv add "dference[pandas]" # if your inputs are pandas DataFrames
76
+ ```
77
+
78
+ `dference` depends on `polars`, `anywidget` and `traitlets`. Inputs can be polars
79
+ (eager or lazy), pandas, pyarrow, or anything implementing the Arrow PyCapsule
80
+ interface or `to_polars()` (DuckDB relations, narwhals, …).
81
+
82
+ ## Quick start
83
+
84
+ ### marimo
85
+
86
+ ```python
87
+ import marimo as mo
88
+ from dference import DataFrameDiff
89
+
90
+ diff = DataFrameDiff(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
91
+ view = mo.ui.anywidget(diff)
92
+ view
93
+ ```
94
+
95
+ In another cell, work with the rows checked in the widget. The cell re-runs
96
+ when the selection changes – and only then; paging, sorting and filtering
97
+ never trigger re-execution:
98
+
99
+ ```python
100
+ view.value["selected_ids"] # reactive dependency
101
+ diff.selected_frame() # the checked rows as a polars DataFrame
102
+ ```
103
+
104
+ ### Jupyter
105
+
106
+ ```python
107
+ from dference import DataFrameDiff
108
+
109
+ DataFrameDiff(crm, erp, key=["region", "customer_id"])
110
+ ```
111
+
112
+ ### Without a UI
113
+
114
+ ```python
115
+ import dference
116
+
117
+ result = dference.compare(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
118
+
119
+ result.summary # Summary(equal=158, mismatch=78, missing_left=10, missing_right=8, ...)
120
+ result.frame() # wide: key, status, differing, "<col> [CRM]", "<col> [ERP]", ...
121
+ result.frame("mismatch")
122
+ result.mismatches() # long: one row per (key, column) that differs
123
+ result.column_stats() # per column: dtypes, mismatches, equal/mismatch share
124
+ ```
125
+
126
+ ## How values are compared
127
+
128
+ - **Exact equality.** There is deliberately no numeric tolerance. If small float
129
+ differences should not count, round both frames first, e.g.
130
+ `df.with_columns(cs.float().round(2))`.
131
+ - **Missing values are equal to each other**: `null == null`, and in float
132
+ columns `NaN` counts as missing too (pandas uses NaN as its null marker, so
133
+ a frame from pandas and one from Arrow would otherwise disagree).
134
+ - **Same column, same dtype.** By default (`strict=True`) a column must have
135
+ the identical dtype on both sides.
136
+ - **`strict=False`** aligns differences that lose nothing – the same values
137
+ stored differently: integers of different width, `Float32` vs. `Float64`,
138
+ decimals of different precision, `Datetime` / `Duration` in different time
139
+ units (same time zone), text vs. `Categorical` / `Enum`, and an all-null
140
+ column vs. anything. `ColumnInfo.compare_dtype` tells you which dtype was
141
+ used.
142
+ - **Everything else is always your decision.** Int vs. float, `Date` vs. `Datetime`,
143
+ naive vs. time-zone aware, different time zones, text vs. numbers, different
144
+ nested types: `compare()` raises one `ValueError` that lists every such column
145
+ (keys included) with a cast to fix it, e.g.
146
+ `right = right.with_columns(pl.col("qty").cast(pl.Int64))`. Note that pandas
147
+ has no date dtype – a `Date` column from pandas arrives as a `Datetime`.
148
+ - **Keys** must be unique on each side; null keys match null keys.
149
+ - **Columns present on one side only** are shown but not compared. Use
150
+ `ignore_columns=[...]` to leave columns out entirely.
151
+
152
+ ## Widget features
153
+
154
+ - Status overview on a shared scale (equal / mismatch / only left / only
155
+ right, plus found vs. not found); filter by status from the menu in the
156
+ status column header.
157
+ - Each side has one colour and one chip – the first letter of its name
158
+ (`left_short=` / `right_short=` to override) – used everywhere: differences
159
+ show both values with equal weight, each marked by its side's chip, and rows
160
+ on one side only carry the chip of the side they exist on.
161
+ - Per-column match rate in the header; click `≠` or `=` to show only rows
162
+ that differ in, or are equal in, that column.
163
+ - Column menu per header: sort and type-aware filters (text, number range,
164
+ date range, boolean, null checks). For compared columns **a row matches if
165
+ the left or the right value matches**; a side that does not exist in a row
166
+ never matches.
167
+ - Invisible characters are made visible: `␣` leading/trailing/repeated spaces,
168
+ `→` tab, `↵` line feed, `⍽` no-break spaces, `ZWSP`, `BOM`, … Cells that
169
+ differ *only* in invisible characters are flagged.
170
+ - Side-by-side detail view with numeric deltas; step through the filtered rows
171
+ across page boundaries.
172
+ - CSV export of all rows matching the current filters.
173
+
174
+ ## Python API
175
+
176
+ | Object | Purpose |
177
+ | --- | --- |
178
+ | `compare(left, right, key, *, left_name, right_name, ignore_columns, strict=True)` | Run a comparison, returns `DiffResult`. |
179
+ | `DiffResult.summary` | `Summary` with counts per status plus `found`, `not_found`, `total`. |
180
+ | `DiffResult.columns` | Tuple of `ColumnInfo` (kind, dtypes, mismatches, compare dtype). |
181
+ | `DiffResult.frame(status=None, *, rows=None)` | Wide result, optionally filtered by status or row ids. |
182
+ | `DiffResult.mismatches()` | Long format of all differing cells (values as text). |
183
+ | `DiffResult.column_stats()` | Match rates per compared column. |
184
+ | `DataFrameDiff(left, right, key, ...)` | The widget; same arguments as `compare` plus `left_short`, `right_short`, `page_size` (default 10). |
185
+ | `DataFrameDiff.from_result(result, ...)` | Widget for an existing `DiffResult`. |
186
+ | `DataFrameDiff.result` | The underlying `DiffResult`. |
187
+ | `DataFrameDiff.selected_ids` | Synced trait with the checked row ids. |
188
+ | `DataFrameDiff.selected_frame()` | Checked rows as a wide frame. |
189
+ | `DataFrameDiff.view_frame()` | Rows matching the widget's current filters (not reactive). |
190
+ | `Status` | `StrEnum` of the four statuses. |
191
+
192
+ ## Performance
193
+
194
+ The comparison is one hash join plus vectorised expressions; the widget keeps
195
+ the ordered row ids of the last few views in a small cache, so paging through
196
+ a view is a slice. Timings from `benchmarks/bench.py` on a **single CPU core**
197
+ (polars parallelises, so more cores are faster):
198
+
199
+ | 5 million rows per side | time |
200
+ | --- | ---: |
201
+ | `compare` (join, difference flags, statistics) | 2.5 s |
202
+ | first page / next page | < 1 ms |
203
+ | filter by status | 19 ms |
204
+ | numeric range filter (left or right) | 45 ms |
205
+ | sort by a column | 0.4 s |
206
+ | full-text search across all columns | 0.4 s |
207
+
208
+ Memory: the joined frame holds both sides once, plus one boolean flag per
209
+ compared column.
210
+
211
+ ## Limitations
212
+
213
+ - The widget needs a running Python kernel; in a static HTML export it shows no
214
+ rows.
215
+ - CSV export sends all matching rows to the browser in one message – filter
216
+ first for very large exports, or use `view_frame().write_csv(...)`.
217
+ - `view_frame()` reflects the filters last sent by the widget but is not a
218
+ reactive value in marimo.
219
+
220
+ ## Development
221
+
222
+ ```bash
223
+ git clone https://github.com/oberbichler/dference && cd dference
224
+ uv sync # creates .venv with all dev tools
225
+
226
+ uv run ruff format . # format
227
+ uv run ruff check . # lint
228
+ uv run ty check # type check
229
+ uv run pytest --cov # tests with coverage
230
+ uv run playwright install chromium # once: browser for the widget tests
231
+ uv run pytest tests/e2e # widget in a real browser (skipped without Chromium)
232
+
233
+ uv run marimo edit examples/demo.py # interactive demo
234
+ uv run python benchmarks/bench.py 1_000_000
235
+ ```
236
+
237
+ The frontend is a single dependency-free ES module
238
+ (`src/dference/static/widget.js`) with its stylesheet next to it – no build step.
239
+ anywidget reloads both on change while you develop.
240
+
241
+ Node.js is only needed for frontend tooling, never at runtime:
242
+
243
+ ```bash
244
+ npm ci # Biome + TypeScript (dev only)
245
+
246
+ npm run fix # format + auto-fix JS, CSS and JSON (Biome)
247
+ npm run check # lint (Biome) and type-check the JSDoc (tsc)
248
+ ```
249
+
250
+ Types live in JSDoc comments and are checked with `tsc --checkJs` in strict
251
+ mode (see `tsconfig.json`); the widget API is typed via `@anywidget/types`.
252
+ Nothing is compiled – the file that ships is the file you edit.
253
+
254
+ ### Releasing
255
+
256
+ The version is derived from the git tag (`v1.2.3`) by
257
+ [hatch-vcs](https://github.com/ofek/hatch-vcs) – there is no version to bump in
258
+ `pyproject.toml` or `uv.lock`. A checkout between releases reports a dev version
259
+ such as `0.1.1.dev3+g94570d5`.
260
+
261
+ For every release:
262
+
263
+ 1. Rename `## [Unreleased]` in `CHANGELOG.md` to `## [x.y.z] - <date>` (add a new
264
+ empty `Unreleased` section above it) and merge that to `main`.
265
+ 2. Run *Actions → Release → Run workflow* on `main` with the version `x.y.z`, or
266
+ `gh workflow run release.yml -f version=x.y.z`.
267
+
268
+ The workflow refuses versions that already exist as tag, release or on PyPI,
269
+ runs the full CI on the commit, builds with that version, publishes to PyPI via
270
+ [trusted publishing](https://docs.pypi.org/trusted-publishers/), and only then
271
+ tags the commit `vx.y.z` and creates the GitHub release with the changelog
272
+ section as notes and the wheel and sdist attached. If a step fails, *Re-run
273
+ failed jobs* continues from there.
274
+
275
+ Once, before the first release: on PyPI add a
276
+ [pending trusted publisher](https://docs.pypi.org/trusted-publishers/creating-a-project-through-oidc/)
277
+ for the project `dference` – owner `oberbichler`, repository `dference`,
278
+ workflow `release.yml`, environment `pypi` – and create the environment `pypi`
279
+ in the GitHub repository settings (optionally with required reviewers).
280
+
281
+ ## License
282
+
283
+ ISC © Thomas Oberbichler
@@ -0,0 +1,249 @@
1
+ <!-- absolute URL: relative image paths do not render on PyPI -->
2
+ <p align="center">
3
+ <img src="https://raw.githubusercontent.com/oberbichler/dference/main/docs/logo.svg" alt="dference" width="480">
4
+ </p>
5
+
6
+ <p align="center">
7
+ <a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/v/dference" alt="PyPI"></a>
8
+ <a href="https://pypi.org/project/dference/"><img src="https://img.shields.io/pypi/pyversions/dference" alt="Python versions"></a>
9
+ <a href="https://pypistats.org/packages/dference"><img src="https://img.shields.io/pypi/dm/dference" alt="Downloads"></a>
10
+ <a href="https://github.com/oberbichler/dference/actions/workflows/ci.yml"><img src="https://github.com/oberbichler/dference/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
11
+ <a href="https://github.com/oberbichler/dference/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-ISC-blue" alt="License: ISC"></a>
12
+ <br>
13
+ <a href="https://github.com/astral-sh/uv"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json" alt="uv"></a>
14
+ <a href="https://github.com/astral-sh/ruff"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json" alt="Ruff"></a>
15
+ <a href="https://github.com/astral-sh/ty"><img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ty/main/assets/badge/v0.json" alt="ty"></a>
16
+ <a href="https://molab.marimo.io/github/oberbichler/dference/blob/main/examples/demo.py"><img src="https://marimo.io/molab-shield.svg" alt="Open the demo in molab"></a>
17
+ </p>
18
+
19
+ **Compare two DataFrames by key and explore every difference – in an interactive
20
+ widget for [marimo](https://marimo.io) and Jupyter, or as plain polars tables.**
21
+
22
+ Rows are matched on a key (one or several columns) and classified as
23
+
24
+ | Status | Meaning |
25
+ | --- | --- |
26
+ | `equal` | key on both sides, all compared values equal |
27
+ | `mismatch` | key on both sides, at least one value differs |
28
+ | `missing_right` | key only in the left frame (shown as `Only in {left_name}`) |
29
+ | `missing_left` | key only in the right frame (shown as `Only in {right_name}`) |
30
+
31
+ The widget is modelled on `marimo.ui.table`: paging, sorting, search, a status
32
+ filter, marimo-style column filters, per-column match rates, a side-by-side
33
+ detail view, invisible-character highlighting and CSV export. It stays fast on
34
+ millions of rows because filtering, sorting and paging run in polars on the
35
+ Python side – the browser only ever receives the visible page.
36
+
37
+ ## Installation
38
+
39
+ ```bash
40
+ uv add dference # or: pip install dference
41
+ uv add "dference[pandas]" # if your inputs are pandas DataFrames
42
+ ```
43
+
44
+ `dference` depends on `polars`, `anywidget` and `traitlets`. Inputs can be polars
45
+ (eager or lazy), pandas, pyarrow, or anything implementing the Arrow PyCapsule
46
+ interface or `to_polars()` (DuckDB relations, narwhals, …).
47
+
48
+ ## Quick start
49
+
50
+ ### marimo
51
+
52
+ ```python
53
+ import marimo as mo
54
+ from dference import DataFrameDiff
55
+
56
+ diff = DataFrameDiff(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
57
+ view = mo.ui.anywidget(diff)
58
+ view
59
+ ```
60
+
61
+ In another cell, work with the rows checked in the widget. The cell re-runs
62
+ when the selection changes – and only then; paging, sorting and filtering
63
+ never trigger re-execution:
64
+
65
+ ```python
66
+ view.value["selected_ids"] # reactive dependency
67
+ diff.selected_frame() # the checked rows as a polars DataFrame
68
+ ```
69
+
70
+ ### Jupyter
71
+
72
+ ```python
73
+ from dference import DataFrameDiff
74
+
75
+ DataFrameDiff(crm, erp, key=["region", "customer_id"])
76
+ ```
77
+
78
+ ### Without a UI
79
+
80
+ ```python
81
+ import dference
82
+
83
+ result = dference.compare(crm, erp, key="customer_id", left_name="CRM", right_name="ERP")
84
+
85
+ result.summary # Summary(equal=158, mismatch=78, missing_left=10, missing_right=8, ...)
86
+ result.frame() # wide: key, status, differing, "<col> [CRM]", "<col> [ERP]", ...
87
+ result.frame("mismatch")
88
+ result.mismatches() # long: one row per (key, column) that differs
89
+ result.column_stats() # per column: dtypes, mismatches, equal/mismatch share
90
+ ```
91
+
92
+ ## How values are compared
93
+
94
+ - **Exact equality.** There is deliberately no numeric tolerance. If small float
95
+ differences should not count, round both frames first, e.g.
96
+ `df.with_columns(cs.float().round(2))`.
97
+ - **Missing values are equal to each other**: `null == null`, and in float
98
+ columns `NaN` counts as missing too (pandas uses NaN as its null marker, so
99
+ a frame from pandas and one from Arrow would otherwise disagree).
100
+ - **Same column, same dtype.** By default (`strict=True`) a column must have
101
+ the identical dtype on both sides.
102
+ - **`strict=False`** aligns differences that lose nothing – the same values
103
+ stored differently: integers of different width, `Float32` vs. `Float64`,
104
+ decimals of different precision, `Datetime` / `Duration` in different time
105
+ units (same time zone), text vs. `Categorical` / `Enum`, and an all-null
106
+ column vs. anything. `ColumnInfo.compare_dtype` tells you which dtype was
107
+ used.
108
+ - **Everything else is always your decision.** Int vs. float, `Date` vs. `Datetime`,
109
+ naive vs. time-zone aware, different time zones, text vs. numbers, different
110
+ nested types: `compare()` raises one `ValueError` that lists every such column
111
+ (keys included) with a cast to fix it, e.g.
112
+ `right = right.with_columns(pl.col("qty").cast(pl.Int64))`. Note that pandas
113
+ has no date dtype – a `Date` column from pandas arrives as a `Datetime`.
114
+ - **Keys** must be unique on each side; null keys match null keys.
115
+ - **Columns present on one side only** are shown but not compared. Use
116
+ `ignore_columns=[...]` to leave columns out entirely.
117
+
118
+ ## Widget features
119
+
120
+ - Status overview on a shared scale (equal / mismatch / only left / only
121
+ right, plus found vs. not found); filter by status from the menu in the
122
+ status column header.
123
+ - Each side has one colour and one chip – the first letter of its name
124
+ (`left_short=` / `right_short=` to override) – used everywhere: differences
125
+ show both values with equal weight, each marked by its side's chip, and rows
126
+ on one side only carry the chip of the side they exist on.
127
+ - Per-column match rate in the header; click `≠` or `=` to show only rows
128
+ that differ in, or are equal in, that column.
129
+ - Column menu per header: sort and type-aware filters (text, number range,
130
+ date range, boolean, null checks). For compared columns **a row matches if
131
+ the left or the right value matches**; a side that does not exist in a row
132
+ never matches.
133
+ - Invisible characters are made visible: `␣` leading/trailing/repeated spaces,
134
+ `→` tab, `↵` line feed, `⍽` no-break spaces, `ZWSP`, `BOM`, … Cells that
135
+ differ *only* in invisible characters are flagged.
136
+ - Side-by-side detail view with numeric deltas; step through the filtered rows
137
+ across page boundaries.
138
+ - CSV export of all rows matching the current filters.
139
+
140
+ ## Python API
141
+
142
+ | Object | Purpose |
143
+ | --- | --- |
144
+ | `compare(left, right, key, *, left_name, right_name, ignore_columns, strict=True)` | Run a comparison, returns `DiffResult`. |
145
+ | `DiffResult.summary` | `Summary` with counts per status plus `found`, `not_found`, `total`. |
146
+ | `DiffResult.columns` | Tuple of `ColumnInfo` (kind, dtypes, mismatches, compare dtype). |
147
+ | `DiffResult.frame(status=None, *, rows=None)` | Wide result, optionally filtered by status or row ids. |
148
+ | `DiffResult.mismatches()` | Long format of all differing cells (values as text). |
149
+ | `DiffResult.column_stats()` | Match rates per compared column. |
150
+ | `DataFrameDiff(left, right, key, ...)` | The widget; same arguments as `compare` plus `left_short`, `right_short`, `page_size` (default 10). |
151
+ | `DataFrameDiff.from_result(result, ...)` | Widget for an existing `DiffResult`. |
152
+ | `DataFrameDiff.result` | The underlying `DiffResult`. |
153
+ | `DataFrameDiff.selected_ids` | Synced trait with the checked row ids. |
154
+ | `DataFrameDiff.selected_frame()` | Checked rows as a wide frame. |
155
+ | `DataFrameDiff.view_frame()` | Rows matching the widget's current filters (not reactive). |
156
+ | `Status` | `StrEnum` of the four statuses. |
157
+
158
+ ## Performance
159
+
160
+ The comparison is one hash join plus vectorised expressions; the widget keeps
161
+ the ordered row ids of the last few views in a small cache, so paging through
162
+ a view is a slice. Timings from `benchmarks/bench.py` on a **single CPU core**
163
+ (polars parallelises, so more cores are faster):
164
+
165
+ | 5 million rows per side | time |
166
+ | --- | ---: |
167
+ | `compare` (join, difference flags, statistics) | 2.5 s |
168
+ | first page / next page | < 1 ms |
169
+ | filter by status | 19 ms |
170
+ | numeric range filter (left or right) | 45 ms |
171
+ | sort by a column | 0.4 s |
172
+ | full-text search across all columns | 0.4 s |
173
+
174
+ Memory: the joined frame holds both sides once, plus one boolean flag per
175
+ compared column.
176
+
177
+ ## Limitations
178
+
179
+ - The widget needs a running Python kernel; in a static HTML export it shows no
180
+ rows.
181
+ - CSV export sends all matching rows to the browser in one message – filter
182
+ first for very large exports, or use `view_frame().write_csv(...)`.
183
+ - `view_frame()` reflects the filters last sent by the widget but is not a
184
+ reactive value in marimo.
185
+
186
+ ## Development
187
+
188
+ ```bash
189
+ git clone https://github.com/oberbichler/dference && cd dference
190
+ uv sync # creates .venv with all dev tools
191
+
192
+ uv run ruff format . # format
193
+ uv run ruff check . # lint
194
+ uv run ty check # type check
195
+ uv run pytest --cov # tests with coverage
196
+ uv run playwright install chromium # once: browser for the widget tests
197
+ uv run pytest tests/e2e # widget in a real browser (skipped without Chromium)
198
+
199
+ uv run marimo edit examples/demo.py # interactive demo
200
+ uv run python benchmarks/bench.py 1_000_000
201
+ ```
202
+
203
+ The frontend is a single dependency-free ES module
204
+ (`src/dference/static/widget.js`) with its stylesheet next to it – no build step.
205
+ anywidget reloads both on change while you develop.
206
+
207
+ Node.js is only needed for frontend tooling, never at runtime:
208
+
209
+ ```bash
210
+ npm ci # Biome + TypeScript (dev only)
211
+
212
+ npm run fix # format + auto-fix JS, CSS and JSON (Biome)
213
+ npm run check # lint (Biome) and type-check the JSDoc (tsc)
214
+ ```
215
+
216
+ Types live in JSDoc comments and are checked with `tsc --checkJs` in strict
217
+ mode (see `tsconfig.json`); the widget API is typed via `@anywidget/types`.
218
+ Nothing is compiled – the file that ships is the file you edit.
219
+
220
+ ### Releasing
221
+
222
+ The version is derived from the git tag (`v1.2.3`) by
223
+ [hatch-vcs](https://github.com/ofek/hatch-vcs) – there is no version to bump in
224
+ `pyproject.toml` or `uv.lock`. A checkout between releases reports a dev version
225
+ such as `0.1.1.dev3+g94570d5`.
226
+
227
+ For every release:
228
+
229
+ 1. Rename `## [Unreleased]` in `CHANGELOG.md` to `## [x.y.z] - <date>` (add a new
230
+ empty `Unreleased` section above it) and merge that to `main`.
231
+ 2. Run *Actions → Release → Run workflow* on `main` with the version `x.y.z`, or
232
+ `gh workflow run release.yml -f version=x.y.z`.
233
+
234
+ The workflow refuses versions that already exist as tag, release or on PyPI,
235
+ runs the full CI on the commit, builds with that version, publishes to PyPI via
236
+ [trusted publishing](https://docs.pypi.org/trusted-publishers/), and only then
237
+ tags the commit `vx.y.z` and creates the GitHub release with the changelog
238
+ section as notes and the wheel and sdist attached. If a step fails, *Re-run
239
+ failed jobs* continues from there.
240
+
241
+ Once, before the first release: on PyPI add a
242
+ [pending trusted publisher](https://docs.pypi.org/trusted-publishers/creating-a-project-through-oidc/)
243
+ for the project `dference` – owner `oberbichler`, repository `dference`,
244
+ workflow `release.yml`, environment `pypi` – and create the environment `pypi`
245
+ in the GitHub repository settings (optionally with required reviewers).
246
+
247
+ ## License
248
+
249
+ ISC © Thomas Oberbichler