databuck-spark-sdk 0.5.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- databuck_spark_sdk-0.5.1/MANIFEST.in +3 -0
- databuck_spark_sdk-0.5.1/PKG-INFO +160 -0
- databuck_spark_sdk-0.5.1/README.md +141 -0
- databuck_spark_sdk-0.5.1/RELEASING.md +42 -0
- databuck_spark_sdk-0.5.1/databuck/__init__.py +8 -0
- databuck_spark_sdk-0.5.1/databuck/__main__.py +7 -0
- databuck_spark_sdk-0.5.1/databuck/agentic_rules.py +346 -0
- databuck_spark_sdk-0.5.1/databuck/lake_rules.py +341 -0
- databuck_spark_sdk-0.5.1/databuck/sdk.py +1328 -0
- databuck_spark_sdk-0.5.1/databuck_spark_sdk.egg-info/PKG-INFO +160 -0
- databuck_spark_sdk-0.5.1/databuck_spark_sdk.egg-info/SOURCES.txt +16 -0
- databuck_spark_sdk-0.5.1/databuck_spark_sdk.egg-info/dependency_links.txt +1 -0
- databuck_spark_sdk-0.5.1/databuck_spark_sdk.egg-info/requires.txt +1 -0
- databuck_spark_sdk-0.5.1/databuck_spark_sdk.egg-info/top_level.txt +1 -0
- databuck_spark_sdk-0.5.1/pyproject.toml +39 -0
- databuck_spark_sdk-0.5.1/setup.cfg +4 -0
- databuck_spark_sdk-0.5.1/tests/test_combined_export.py +83 -0
- databuck_spark_sdk-0.5.1/tests/test_jar_download.py +115 -0
|
@@ -0,0 +1,160 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: databuck-spark-sdk
|
|
3
|
+
Version: 0.5.1
|
|
4
|
+
Summary: DataBuck data quality SDK for PySpark and Databricks
|
|
5
|
+
Author: DataBuck
|
|
6
|
+
Project-URL: Repository, https://github.com/FirstEigen-Labs/databuck_sdk
|
|
7
|
+
Keywords: data quality,databricks,pyspark,data profiling
|
|
8
|
+
Classifier: Intended Audience :: Developers
|
|
9
|
+
Classifier: Operating System :: OS Independent
|
|
10
|
+
Classifier: Programming Language :: Python :: 3
|
|
11
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
15
|
+
Classifier: Topic :: Database
|
|
16
|
+
Requires-Python: >=3.10
|
|
17
|
+
Description-Content-Type: text/markdown
|
|
18
|
+
Requires-Dist: google-genai>=1.0.0
|
|
19
|
+
|
|
20
|
+
# databuck-spark-sdk
|
|
21
|
+
|
|
22
|
+
Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.
|
|
23
|
+
|
|
24
|
+
## Install
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
python -m pip install databuck-spark-sdk
|
|
28
|
+
python -c "import databuck"
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
Importing `databuck` downloads `databuck-spark-sdk.jar` from the DataBuck S3
|
|
32
|
+
URL to `databuck/jars/` in the installed Python package, with progress output.
|
|
33
|
+
The downloaded path is set in `DATABUCK_SPARK_SDK_JAR` for the current Python
|
|
34
|
+
process. A later import reuses the existing JAR.
|
|
35
|
+
The JAR is not included in the wheel because it exceeds PyPI's
|
|
36
|
+
default per-file upload limit. An internet connection and write access to the
|
|
37
|
+
installed package directory are required for the first download.
|
|
38
|
+
|
|
39
|
+
`python -m databuck` is an explicit download command that also prints the
|
|
40
|
+
local path. Set `DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0` before import to defer the
|
|
41
|
+
download when offline. `DataBuck.jar_path()` returns the path after download.
|
|
42
|
+
Downloading the JAR does not install it as a Databricks compute library.
|
|
43
|
+
|
|
44
|
+
To use an existing JAR or another writable destination, set
|
|
45
|
+
`DATABUCK_SPARK_SDK_JAR` to that file path. If the download URL changes, set
|
|
46
|
+
`DATABUCK_SPARK_SDK_JAR_URL` to the new HTTPS URL before importing or running
|
|
47
|
+
`python -m databuck`. The download is checked against
|
|
48
|
+
the published JAR's SHA-256 hash. If the JAR content changes, publish a new
|
|
49
|
+
Python package version with its new hash, or set
|
|
50
|
+
`DATABUCK_SPARK_SDK_JAR_SHA256` to the expected hash.
|
|
51
|
+
|
|
52
|
+
For a Databricks notebook, install the Python distribution as a library and
|
|
53
|
+
make the JAR available on the cluster. If the package directory is read-only,
|
|
54
|
+
configure `DATABUCK_SPARK_SDK_JAR` to a writable driver path before importing.
|
|
55
|
+
|
|
56
|
+
## Usage
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
from pyspark.sql import SparkSession
|
|
60
|
+
import os
|
|
61
|
+
|
|
62
|
+
from databuck import DataBuck
|
|
63
|
+
|
|
64
|
+
spark = (
|
|
65
|
+
SparkSession.builder
|
|
66
|
+
.config("spark.jars", DataBuck.jar_path())
|
|
67
|
+
.getOrCreate()
|
|
68
|
+
)
|
|
69
|
+
|
|
70
|
+
df = spark.read.csv("customers.csv", header=True)
|
|
71
|
+
|
|
72
|
+
print(DataBuck.count(df))
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
## Discover and export rules for Databricks
|
|
76
|
+
|
|
77
|
+
```python
|
|
78
|
+
json_path = DataBuck.discover_and_export(
|
|
79
|
+
df,
|
|
80
|
+
"/Volumes/catalog/schema/volume/expectations.json",
|
|
81
|
+
context={
|
|
82
|
+
"business_context": (
|
|
83
|
+
"Missing customer email should be monitored. "
|
|
84
|
+
"Orders without an order ID must be discarded. "
|
|
85
|
+
"An unknown currency must stop publication."
|
|
86
|
+
),
|
|
87
|
+
"gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
|
|
88
|
+
},
|
|
89
|
+
)
|
|
90
|
+
print(json_path)
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
The single call runs profile discovery and context-aware BuckGPT discovery,
|
|
94
|
+
then writes both sets of row expectations into one JSON file. Context-aware
|
|
95
|
+
expectations are named `BuckGPT_Rule_001`, `BuckGPT_Rule_002`, etc. Their
|
|
96
|
+
invalid-row SQL queries are converted to Lakeflow pass conditions. Only
|
|
97
|
+
audited, passed BuckGPT rules in the required
|
|
98
|
+
`SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition>` shape can be
|
|
99
|
+
exported; an unsupported query fails the export.
|
|
100
|
+
|
|
101
|
+
To export only automatic profiling rules, omit `context`:
|
|
102
|
+
|
|
103
|
+
```python
|
|
104
|
+
json_path = DataBuck.discover_and_export(
|
|
105
|
+
df, "/Volumes/catalog/schema/volume/expectations.json"
|
|
106
|
+
)
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
Without context, BuckGPT is not called and all exported rules use `warn`.
|
|
110
|
+
|
|
111
|
+
The JSON contains three dictionaries: `warn`, `drop`, and `fail`. Gemini
|
|
112
|
+
classifies every exported expectation using the business context, DataFrame
|
|
113
|
+
schema, and up to five sample rows. This sends those inputs to Gemini. If an
|
|
114
|
+
LLM response is incomplete or invalid, export fails instead of assigning a
|
|
115
|
+
destructive action.
|
|
116
|
+
In a Lakeflow pipeline, use the dictionaries with `dp.expect_all`,
|
|
117
|
+
`dp.expect_all_or_drop`, and `dp.expect_all_or_fail`, respectively.
|
|
118
|
+
|
|
119
|
+
For example, in the Lakeflow pipeline source file:
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
import json
|
|
123
|
+
from pyspark import pipelines as dp
|
|
124
|
+
|
|
125
|
+
with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
|
|
126
|
+
expectations = json.load(stream)
|
|
127
|
+
|
|
128
|
+
@dp.table
|
|
129
|
+
@dp.expect_all(expectations["warn"])
|
|
130
|
+
@dp.expect_all_or_drop(expectations["drop"])
|
|
131
|
+
@dp.expect_all_or_fail(expectations["fail"])
|
|
132
|
+
def customers_checked():
|
|
133
|
+
return spark.read.table("catalog.schema.customers_source")
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
The pipeline must read a DataFrame with the columns used by the exported
|
|
137
|
+
expectations. Regenerate the JSON when the source schema or business policy
|
|
138
|
+
changes, then refresh the pipeline.
|
|
139
|
+
|
|
140
|
+
```json
|
|
141
|
+
{
|
|
142
|
+
"warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
|
|
143
|
+
"drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
|
|
144
|
+
"fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
|
|
145
|
+
}
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
The example shows the file shape; the actual actions come from the supplied
|
|
149
|
+
business context. `DataBuck.discover_rules(df)` and
|
|
150
|
+
`DataBuck.discover(df, context)` are still available separately. The new
|
|
151
|
+
`DataBuck.discover_and_export(...)` call combines them.
|
|
152
|
+
|
|
153
|
+
Each action dictionary maps expectation names to Spark SQL pass conditions,
|
|
154
|
+
such as `{"not_null_subscriber_id": "`subscriber_id` IS NOT NULL"}`. Null, pattern,
|
|
155
|
+
and length rules are converted from their profiling metadata, including rules
|
|
156
|
+
from older SDK JARs whose `expression` field is empty. Dataset-level rules
|
|
157
|
+
and catalog rules without a row condition remain in `rules` but are omitted
|
|
158
|
+
from the file. Null rules are exported only when their threshold is 0%.
|
|
159
|
+
The export does not encode aggregate failure thresholds. Profiling pattern
|
|
160
|
+
`A` maps to `[A-Z]`, and `#` maps to `[0-9]`.
|
|
@@ -0,0 +1,141 @@
|
|
|
1
|
+
# databuck-spark-sdk
|
|
2
|
+
|
|
3
|
+
Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.
|
|
4
|
+
|
|
5
|
+
## Install
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
python -m pip install databuck-spark-sdk
|
|
9
|
+
python -c "import databuck"
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
Importing `databuck` downloads `databuck-spark-sdk.jar` from the DataBuck S3
|
|
13
|
+
URL to `databuck/jars/` in the installed Python package, with progress output.
|
|
14
|
+
The downloaded path is set in `DATABUCK_SPARK_SDK_JAR` for the current Python
|
|
15
|
+
process. A later import reuses the existing JAR.
|
|
16
|
+
The JAR is not included in the wheel because it exceeds PyPI's
|
|
17
|
+
default per-file upload limit. An internet connection and write access to the
|
|
18
|
+
installed package directory are required for the first download.
|
|
19
|
+
|
|
20
|
+
`python -m databuck` is an explicit download command that also prints the
|
|
21
|
+
local path. Set `DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0` before import to defer the
|
|
22
|
+
download when offline. `DataBuck.jar_path()` returns the path after download.
|
|
23
|
+
Downloading the JAR does not install it as a Databricks compute library.
|
|
24
|
+
|
|
25
|
+
To use an existing JAR or another writable destination, set
|
|
26
|
+
`DATABUCK_SPARK_SDK_JAR` to that file path. If the download URL changes, set
|
|
27
|
+
`DATABUCK_SPARK_SDK_JAR_URL` to the new HTTPS URL before importing or running
|
|
28
|
+
`python -m databuck`. The download is checked against
|
|
29
|
+
the published JAR's SHA-256 hash. If the JAR content changes, publish a new
|
|
30
|
+
Python package version with its new hash, or set
|
|
31
|
+
`DATABUCK_SPARK_SDK_JAR_SHA256` to the expected hash.
|
|
32
|
+
|
|
33
|
+
For a Databricks notebook, install the Python distribution as a library and
|
|
34
|
+
make the JAR available on the cluster. If the package directory is read-only,
|
|
35
|
+
configure `DATABUCK_SPARK_SDK_JAR` to a writable driver path before importing.
|
|
36
|
+
|
|
37
|
+
## Usage
|
|
38
|
+
|
|
39
|
+
```python
|
|
40
|
+
from pyspark.sql import SparkSession
|
|
41
|
+
import os
|
|
42
|
+
|
|
43
|
+
from databuck import DataBuck
|
|
44
|
+
|
|
45
|
+
spark = (
|
|
46
|
+
SparkSession.builder
|
|
47
|
+
.config("spark.jars", DataBuck.jar_path())
|
|
48
|
+
.getOrCreate()
|
|
49
|
+
)
|
|
50
|
+
|
|
51
|
+
df = spark.read.csv("customers.csv", header=True)
|
|
52
|
+
|
|
53
|
+
print(DataBuck.count(df))
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
## Discover and export rules for Databricks
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
json_path = DataBuck.discover_and_export(
|
|
60
|
+
df,
|
|
61
|
+
"/Volumes/catalog/schema/volume/expectations.json",
|
|
62
|
+
context={
|
|
63
|
+
"business_context": (
|
|
64
|
+
"Missing customer email should be monitored. "
|
|
65
|
+
"Orders without an order ID must be discarded. "
|
|
66
|
+
"An unknown currency must stop publication."
|
|
67
|
+
),
|
|
68
|
+
"gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
|
|
69
|
+
},
|
|
70
|
+
)
|
|
71
|
+
print(json_path)
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
The single call runs profile discovery and context-aware BuckGPT discovery,
|
|
75
|
+
then writes both sets of row expectations into one JSON file. Context-aware
|
|
76
|
+
expectations are named `BuckGPT_Rule_001`, `BuckGPT_Rule_002`, etc. Their
|
|
77
|
+
invalid-row SQL queries are converted to Lakeflow pass conditions. Only
|
|
78
|
+
audited, passed BuckGPT rules in the required
|
|
79
|
+
`SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition>` shape can be
|
|
80
|
+
exported; an unsupported query fails the export.
|
|
81
|
+
|
|
82
|
+
To export only automatic profiling rules, omit `context`:
|
|
83
|
+
|
|
84
|
+
```python
|
|
85
|
+
json_path = DataBuck.discover_and_export(
|
|
86
|
+
df, "/Volumes/catalog/schema/volume/expectations.json"
|
|
87
|
+
)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
Without context, BuckGPT is not called and all exported rules use `warn`.
|
|
91
|
+
|
|
92
|
+
The JSON contains three dictionaries: `warn`, `drop`, and `fail`. Gemini
|
|
93
|
+
classifies every exported expectation using the business context, DataFrame
|
|
94
|
+
schema, and up to five sample rows. This sends those inputs to Gemini. If an
|
|
95
|
+
LLM response is incomplete or invalid, export fails instead of assigning a
|
|
96
|
+
destructive action.
|
|
97
|
+
In a Lakeflow pipeline, use the dictionaries with `dp.expect_all`,
|
|
98
|
+
`dp.expect_all_or_drop`, and `dp.expect_all_or_fail`, respectively.
|
|
99
|
+
|
|
100
|
+
For example, in the Lakeflow pipeline source file:
|
|
101
|
+
|
|
102
|
+
```python
|
|
103
|
+
import json
|
|
104
|
+
from pyspark import pipelines as dp
|
|
105
|
+
|
|
106
|
+
with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
|
|
107
|
+
expectations = json.load(stream)
|
|
108
|
+
|
|
109
|
+
@dp.table
|
|
110
|
+
@dp.expect_all(expectations["warn"])
|
|
111
|
+
@dp.expect_all_or_drop(expectations["drop"])
|
|
112
|
+
@dp.expect_all_or_fail(expectations["fail"])
|
|
113
|
+
def customers_checked():
|
|
114
|
+
return spark.read.table("catalog.schema.customers_source")
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The pipeline must read a DataFrame with the columns used by the exported
|
|
118
|
+
expectations. Regenerate the JSON when the source schema or business policy
|
|
119
|
+
changes, then refresh the pipeline.
|
|
120
|
+
|
|
121
|
+
```json
|
|
122
|
+
{
|
|
123
|
+
"warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
|
|
124
|
+
"drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
|
|
125
|
+
"fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
|
|
126
|
+
}
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
The example shows the file shape; the actual actions come from the supplied
|
|
130
|
+
business context. `DataBuck.discover_rules(df)` and
|
|
131
|
+
`DataBuck.discover(df, context)` are still available separately. The new
|
|
132
|
+
`DataBuck.discover_and_export(...)` call combines them.
|
|
133
|
+
|
|
134
|
+
Each action dictionary maps expectation names to Spark SQL pass conditions,
|
|
135
|
+
such as `{"not_null_subscriber_id": "`subscriber_id` IS NOT NULL"}`. Null, pattern,
|
|
136
|
+
and length rules are converted from their profiling metadata, including rules
|
|
137
|
+
from older SDK JARs whose `expression` field is empty. Dataset-level rules
|
|
138
|
+
and catalog rules without a row condition remain in `rules` but are omitted
|
|
139
|
+
from the file. Null rules are exported only when their threshold is 0%.
|
|
140
|
+
The export does not encode aggregate failure thresholds. Profiling pattern
|
|
141
|
+
`A` maps to `[A-Z]`, and `#` maps to `[0-9]`.
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# Publish with GitHub Actions trusted publishing
|
|
2
|
+
|
|
3
|
+
The publishing workflow is `.github/workflows/publish.yml` at the repository
|
|
4
|
+
root. It runs only when manually started on `main`. It tests and builds the
|
|
5
|
+
Python package from `databuck-python/databuck-spark/`, checks the distributions,
|
|
6
|
+
and publishes them to PyPI without a manually created API token.
|
|
7
|
+
|
|
8
|
+
For the first release, add a pending trusted publisher on PyPI with:
|
|
9
|
+
|
|
10
|
+
- PyPI project name: `databuck-spark-sdk`
|
|
11
|
+
- Owner: `FirstEigen-Labs`
|
|
12
|
+
- Repository name: `databuck_sdk`
|
|
13
|
+
- Workflow name: `publish.yml`
|
|
14
|
+
- Environment name: leave blank
|
|
15
|
+
|
|
16
|
+
Commit and push the workflow to `main` before running it. Then in GitHub, open
|
|
17
|
+
**Actions → Publish databuck-spark-sdk to PyPI → Run workflow**, select `main`,
|
|
18
|
+
and start the run. Check its result and the new release on PyPI.
|
|
19
|
+
|
|
20
|
+
Each PyPI release needs a new version in `pyproject.toml`; release files for an
|
|
21
|
+
existing version cannot be overwritten. Review the version and source changes
|
|
22
|
+
before starting the workflow.
|
|
23
|
+
|
|
24
|
+
## Optional local build check
|
|
25
|
+
|
|
26
|
+
Run these commands from `databuck-spark/` after setting the intended version
|
|
27
|
+
in `pyproject.toml`:
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
python -m pip install --upgrade build twine
|
|
31
|
+
python -m build --outdir dist-release-NEW_VERSION
|
|
32
|
+
python -m twine check dist-release-NEW_VERSION/*
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
Inspect both generated files (`.whl` and `.tar.gz`) before upload. The wheel
|
|
36
|
+
must contain the `databuck` package and its metadata. The JAR is downloaded
|
|
37
|
+
when `databuck` is imported or by `python -m databuck` after installation. Check that its
|
|
38
|
+
public S3 URL works without credentials before publishing. Do not stage the
|
|
39
|
+
large JAR in the Python distribution.
|
|
40
|
+
|
|
41
|
+
Use a fresh output directory for each build. On Windows, rebuilding into a
|
|
42
|
+
directory that already contains the same archive can fail with `WinError 5`.
|