databuck-spark-sdk 0.5.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,3 @@
1
+ include README.md
2
+ include RELEASING.md
3
+ global-exclude *.jar
@@ -0,0 +1,160 @@
1
+ Metadata-Version: 2.4
2
+ Name: databuck-spark-sdk
3
+ Version: 0.5.1
4
+ Summary: DataBuck data quality SDK for PySpark and Databricks
5
+ Author: DataBuck
6
+ Project-URL: Repository, https://github.com/FirstEigen-Labs/databuck_sdk
7
+ Keywords: data quality,databricks,pyspark,data profiling
8
+ Classifier: Intended Audience :: Developers
9
+ Classifier: Operating System :: OS Independent
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Programming Language :: Python :: 3.10
12
+ Classifier: Programming Language :: Python :: 3.11
13
+ Classifier: Programming Language :: Python :: 3.12
14
+ Classifier: Programming Language :: Python :: 3.13
15
+ Classifier: Topic :: Database
16
+ Requires-Python: >=3.10
17
+ Description-Content-Type: text/markdown
18
+ Requires-Dist: google-genai>=1.0.0
19
+
20
+ # databuck-spark-sdk
21
+
22
+ Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.
23
+
24
+ ## Install
25
+
26
+ ```bash
27
+ python -m pip install databuck-spark-sdk
28
+ python -c "import databuck"
29
+ ```
30
+
31
+ Importing `databuck` downloads `databuck-spark-sdk.jar` from the DataBuck S3
32
+ URL to `databuck/jars/` in the installed Python package, with progress output.
33
+ The downloaded path is set in `DATABUCK_SPARK_SDK_JAR` for the current Python
34
+ process. A later import reuses the existing JAR.
35
+ The JAR is not included in the wheel because it exceeds PyPI's
36
+ default per-file upload limit. An internet connection and write access to the
37
+ installed package directory are required for the first download.
38
+
39
+ `python -m databuck` is an explicit download command that also prints the
40
+ local path. Set `DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0` before import to defer the
41
+ download when offline. `DataBuck.jar_path()` returns the path after download.
42
+ Downloading the JAR does not install it as a Databricks compute library.
43
+
44
+ To use an existing JAR or another writable destination, set
45
+ `DATABUCK_SPARK_SDK_JAR` to that file path. If the download URL changes, set
46
+ `DATABUCK_SPARK_SDK_JAR_URL` to the new HTTPS URL before importing or running
47
+ `python -m databuck`. The download is checked against
48
+ the published JAR's SHA-256 hash. If the JAR content changes, publish a new
49
+ Python package version with its new hash, or set
50
+ `DATABUCK_SPARK_SDK_JAR_SHA256` to the expected hash.
51
+
52
+ For a Databricks notebook, install the Python distribution as a library and
53
+ make the JAR available on the cluster. If the package directory is read-only,
54
+ configure `DATABUCK_SPARK_SDK_JAR` to a writable driver path before importing.
55
+
56
+ ## Usage
57
+
58
+ ```python
59
+ from pyspark.sql import SparkSession
60
+ import os
61
+
62
+ from databuck import DataBuck
63
+
64
+ spark = (
65
+ SparkSession.builder
66
+ .config("spark.jars", DataBuck.jar_path())
67
+ .getOrCreate()
68
+ )
69
+
70
+ df = spark.read.csv("customers.csv", header=True)
71
+
72
+ print(DataBuck.count(df))
73
+ ```
74
+
75
+ ## Discover and export rules for Databricks
76
+
77
+ ```python
78
+ json_path = DataBuck.discover_and_export(
79
+ df,
80
+ "/Volumes/catalog/schema/volume/expectations.json",
81
+ context={
82
+ "business_context": (
83
+ "Missing customer email should be monitored. "
84
+ "Orders without an order ID must be discarded. "
85
+ "An unknown currency must stop publication."
86
+ ),
87
+ "gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
88
+ },
89
+ )
90
+ print(json_path)
91
+ ```
92
+
93
+ The single call runs profile discovery and context-aware BuckGPT discovery,
94
+ then writes both sets of row expectations into one JSON file. Context-aware
95
+ expectations are named `BuckGPT_Rule_001`, `BuckGPT_Rule_002`, etc. Their
96
+ invalid-row SQL queries are converted to Lakeflow pass conditions. Only
97
+ audited, passed BuckGPT rules in the required
98
+ `SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition>` shape can be
99
+ exported; an unsupported query fails the export.
100
+
101
+ To export only automatic profiling rules, omit `context`:
102
+
103
+ ```python
104
+ json_path = DataBuck.discover_and_export(
105
+ df, "/Volumes/catalog/schema/volume/expectations.json"
106
+ )
107
+ ```
108
+
109
+ Without context, BuckGPT is not called and all exported rules use `warn`.
110
+
111
+ The JSON contains three dictionaries: `warn`, `drop`, and `fail`. Gemini
112
+ classifies every exported expectation using the business context, DataFrame
113
+ schema, and up to five sample rows. This sends those inputs to Gemini. If an
114
+ LLM response is incomplete or invalid, export fails instead of assigning a
115
+ destructive action.
116
+ In a Lakeflow pipeline, use the dictionaries with `dp.expect_all`,
117
+ `dp.expect_all_or_drop`, and `dp.expect_all_or_fail`, respectively.
118
+
119
+ For example, in the Lakeflow pipeline source file:
120
+
121
+ ```python
122
+ import json
123
+ from pyspark import pipelines as dp
124
+
125
+ with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
126
+ expectations = json.load(stream)
127
+
128
+ @dp.table
129
+ @dp.expect_all(expectations["warn"])
130
+ @dp.expect_all_or_drop(expectations["drop"])
131
+ @dp.expect_all_or_fail(expectations["fail"])
132
+ def customers_checked():
133
+ return spark.read.table("catalog.schema.customers_source")
134
+ ```
135
+
136
+ The pipeline must read a DataFrame with the columns used by the exported
137
+ expectations. Regenerate the JSON when the source schema or business policy
138
+ changes, then refresh the pipeline.
139
+
140
+ ```json
141
+ {
142
+ "warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
143
+ "drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
144
+ "fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
145
+ }
146
+ ```
147
+
148
+ The example shows the file shape; the actual actions come from the supplied
149
+ business context. `DataBuck.discover_rules(df)` and
150
+ `DataBuck.discover(df, context)` are still available separately. The new
151
+ `DataBuck.discover_and_export(...)` call combines them.
152
+
153
+ Each action dictionary maps expectation names to Spark SQL pass conditions,
154
+ such as `{"not_null_subscriber_id": "`subscriber_id` IS NOT NULL"}`. Null, pattern,
155
+ and length rules are converted from their profiling metadata, including rules
156
+ from older SDK JARs whose `expression` field is empty. Dataset-level rules
157
+ and catalog rules without a row condition remain in `rules` but are omitted
158
+ from the file. Null rules are exported only when their threshold is 0%.
159
+ The export does not encode aggregate failure thresholds. Profiling pattern
160
+ `A` maps to `[A-Z]`, and `#` maps to `[0-9]`.
@@ -0,0 +1,141 @@
1
+ # databuck-spark-sdk
2
+
3
+ Python wrapper for the DataBuck Spark SDK, used from PySpark and Databricks.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ python -m pip install databuck-spark-sdk
9
+ python -c "import databuck"
10
+ ```
11
+
12
+ Importing `databuck` downloads `databuck-spark-sdk.jar` from the DataBuck S3
13
+ URL to `databuck/jars/` in the installed Python package, with progress output.
14
+ The downloaded path is set in `DATABUCK_SPARK_SDK_JAR` for the current Python
15
+ process. A later import reuses the existing JAR.
16
+ The JAR is not included in the wheel because it exceeds PyPI's
17
+ default per-file upload limit. An internet connection and write access to the
18
+ installed package directory are required for the first download.
19
+
20
+ `python -m databuck` is an explicit download command that also prints the
21
+ local path. Set `DATABUCK_SPARK_SDK_AUTO_DOWNLOAD=0` before import to defer the
22
+ download when offline. `DataBuck.jar_path()` returns the path after download.
23
+ Downloading the JAR does not install it as a Databricks compute library.
24
+
25
+ To use an existing JAR or another writable destination, set
26
+ `DATABUCK_SPARK_SDK_JAR` to that file path. If the download URL changes, set
27
+ `DATABUCK_SPARK_SDK_JAR_URL` to the new HTTPS URL before importing or running
28
+ `python -m databuck`. The download is checked against
29
+ the published JAR's SHA-256 hash. If the JAR content changes, publish a new
30
+ Python package version with its new hash, or set
31
+ `DATABUCK_SPARK_SDK_JAR_SHA256` to the expected hash.
32
+
33
+ For a Databricks notebook, install the Python distribution as a library and
34
+ make the JAR available on the cluster. If the package directory is read-only,
35
+ configure `DATABUCK_SPARK_SDK_JAR` to a writable driver path before importing.
36
+
37
+ ## Usage
38
+
39
+ ```python
40
+ from pyspark.sql import SparkSession
41
+ import os
42
+
43
+ from databuck import DataBuck
44
+
45
+ spark = (
46
+ SparkSession.builder
47
+ .config("spark.jars", DataBuck.jar_path())
48
+ .getOrCreate()
49
+ )
50
+
51
+ df = spark.read.csv("customers.csv", header=True)
52
+
53
+ print(DataBuck.count(df))
54
+ ```
55
+
56
+ ## Discover and export rules for Databricks
57
+
58
+ ```python
59
+ json_path = DataBuck.discover_and_export(
60
+ df,
61
+ "/Volumes/catalog/schema/volume/expectations.json",
62
+ context={
63
+ "business_context": (
64
+ "Missing customer email should be monitored. "
65
+ "Orders without an order ID must be discarded. "
66
+ "An unknown currency must stop publication."
67
+ ),
68
+ "gemini_api_key": dbutils.secrets.get(scope="databuck", key="gemini-api-key"),
69
+ },
70
+ )
71
+ print(json_path)
72
+ ```
73
+
74
+ The single call runs profile discovery and context-aware BuckGPT discovery,
75
+ then writes both sets of row expectations into one JSON file. Context-aware
76
+ expectations are named `BuckGPT_Rule_001`, `BuckGPT_Rule_002`, etc. Their
77
+ invalid-row SQL queries are converted to Lakeflow pass conditions. Only
78
+ audited, passed BuckGPT rules in the required
79
+ `SELECT * FROM {{DATAFRAME}} WHERE <invalid-row condition>` shape can be
80
+ exported; an unsupported query fails the export.
81
+
82
+ To export only automatic profiling rules, omit `context`:
83
+
84
+ ```python
85
+ json_path = DataBuck.discover_and_export(
86
+ df, "/Volumes/catalog/schema/volume/expectations.json"
87
+ )
88
+ ```
89
+
90
+ Without context, BuckGPT is not called and all exported rules use `warn`.
91
+
92
+ The JSON contains three dictionaries: `warn`, `drop`, and `fail`. Gemini
93
+ classifies every exported expectation using the business context, DataFrame
94
+ schema, and up to five sample rows. This sends those inputs to Gemini. If an
95
+ LLM response is incomplete or invalid, export fails instead of assigning a
96
+ destructive action.
97
+ In a Lakeflow pipeline, use the dictionaries with `dp.expect_all`,
98
+ `dp.expect_all_or_drop`, and `dp.expect_all_or_fail`, respectively.
99
+
100
+ For example, in the Lakeflow pipeline source file:
101
+
102
+ ```python
103
+ import json
104
+ from pyspark import pipelines as dp
105
+
106
+ with open("/Volumes/catalog/schema/volume/expectations.json", encoding="utf-8") as stream:
107
+ expectations = json.load(stream)
108
+
109
+ @dp.table
110
+ @dp.expect_all(expectations["warn"])
111
+ @dp.expect_all_or_drop(expectations["drop"])
112
+ @dp.expect_all_or_fail(expectations["fail"])
113
+ def customers_checked():
114
+ return spark.read.table("catalog.schema.customers_source")
115
+ ```
116
+
117
+ The pipeline must read a DataFrame with the columns used by the exported
118
+ expectations. Regenerate the JSON when the source schema or business policy
119
+ changes, then refresh the pipeline.
120
+
121
+ ```json
122
+ {
123
+ "warn": {"not_null_customer_email": "`customer_email` IS NOT NULL"},
124
+ "drop": {"not_null_order_id": "`order_id` IS NOT NULL"},
125
+ "fail": {"valid_pattern_currency_code": "`currency_code` IS NULL OR CAST(`currency_code` AS STRING) RLIKE '(?:^[A-Z][A-Z][A-Z]$)'"}
126
+ }
127
+ ```
128
+
129
+ The example shows the file shape; the actual actions come from the supplied
130
+ business context. `DataBuck.discover_rules(df)` and
131
+ `DataBuck.discover(df, context)` are still available separately. The new
132
+ `DataBuck.discover_and_export(...)` call combines them.
133
+
134
+ Each action dictionary maps expectation names to Spark SQL pass conditions,
135
+ such as `{"not_null_subscriber_id": "`subscriber_id` IS NOT NULL"}`. Null, pattern,
136
+ and length rules are converted from their profiling metadata, including rules
137
+ from older SDK JARs whose `expression` field is empty. Dataset-level rules
138
+ and catalog rules without a row condition remain in `rules` but are omitted
139
+ from the file. Null rules are exported only when their threshold is 0%.
140
+ The export does not encode aggregate failure thresholds. Profiling pattern
141
+ `A` maps to `[A-Z]`, and `#` maps to `[0-9]`.
@@ -0,0 +1,42 @@
1
+ # Publish with GitHub Actions trusted publishing
2
+
3
+ The publishing workflow is `.github/workflows/publish.yml` at the repository
4
+ root. It runs only when manually started on `main`. It tests and builds the
5
+ Python package from `databuck-python/databuck-spark/`, checks the distributions,
6
+ and publishes them to PyPI without a manually created API token.
7
+
8
+ For the first release, add a pending trusted publisher on PyPI with:
9
+
10
+ - PyPI project name: `databuck-spark-sdk`
11
+ - Owner: `FirstEigen-Labs`
12
+ - Repository name: `databuck_sdk`
13
+ - Workflow name: `publish.yml`
14
+ - Environment name: leave blank
15
+
16
+ Commit and push the workflow to `main` before running it. Then in GitHub, open
17
+ **Actions → Publish databuck-spark-sdk to PyPI → Run workflow**, select `main`,
18
+ and start the run. Check its result and the new release on PyPI.
19
+
20
+ Each PyPI release needs a new version in `pyproject.toml`; release files for an
21
+ existing version cannot be overwritten. Review the version and source changes
22
+ before starting the workflow.
23
+
24
+ ## Optional local build check
25
+
26
+ Run these commands from `databuck-spark/` after setting the intended version
27
+ in `pyproject.toml`:
28
+
29
+ ```bash
30
+ python -m pip install --upgrade build twine
31
+ python -m build --outdir dist-release-NEW_VERSION
32
+ python -m twine check dist-release-NEW_VERSION/*
33
+ ```
34
+
35
+ Inspect both generated files (`.whl` and `.tar.gz`) before upload. The wheel
36
+ must contain the `databuck` package and its metadata. The JAR is downloaded
37
+ when `databuck` is imported or by `python -m databuck` after installation. Check that its
38
+ public S3 URL works without credentials before publishing. Do not stage the
39
+ large JAR in the Python distribution.
40
+
41
+ Use a fresh output directory for each build. On Windows, rebuilding into a
42
+ directory that already contains the same archive can fail with `WinError 5`.
@@ -0,0 +1,8 @@
1
+ import os
2
+
3
+ from .sdk import DataBuck
4
+
5
+ if os.environ.get("DATABUCK_SPARK_SDK_AUTO_DOWNLOAD", "1").lower() not in ("0", "false", "no"):
6
+ os.environ["DATABUCK_SPARK_SDK_JAR"] = DataBuck.download_jar()
7
+
8
+ __all__ = ["DataBuck"]
@@ -0,0 +1,7 @@
1
+ """Download the DataBuck Spark SDK JAR after installing the Python package."""
2
+
3
+ from .sdk import DataBuck
4
+
5
+
6
+ if __name__ == "__main__":
7
+ print(DataBuck.download_jar())