ml-gearbox 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ml_gearbox-0.1.0/PKG-INFO +270 -0
- ml_gearbox-0.1.0/README.md +254 -0
- ml_gearbox-0.1.0/pyproject.toml +29 -0
- ml_gearbox-0.1.0/setup.cfg +4 -0
- ml_gearbox-0.1.0/src/ml_gearbox/__init__.py +0 -0
- ml_gearbox-0.1.0/src/ml_gearbox/audit.py +458 -0
- ml_gearbox-0.1.0/src/ml_gearbox/eda.py +219 -0
- ml_gearbox-0.1.0/src/ml_gearbox/splitting.py +0 -0
- ml_gearbox-0.1.0/src/ml_gearbox/utils.py +2 -0
- ml_gearbox-0.1.0/src/ml_gearbox.egg-info/PKG-INFO +270 -0
- ml_gearbox-0.1.0/src/ml_gearbox.egg-info/SOURCES.txt +12 -0
- ml_gearbox-0.1.0/src/ml_gearbox.egg-info/dependency_links.txt +1 -0
- ml_gearbox-0.1.0/src/ml_gearbox.egg-info/requires.txt +7 -0
- ml_gearbox-0.1.0/src/ml_gearbox.egg-info/top_level.txt +1 -0
|
@@ -0,0 +1,270 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: ml-gearbox
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Reusable Machine Learning Toolkit
|
|
5
|
+
Author: Andrew Wamala Kyanjo
|
|
6
|
+
Project-URL: Homepage, https://github.com/AndrewKyanjo/ml-toolkit
|
|
7
|
+
Project-URL: Issues, https://github.com/AndrewKyanjo/ml-toolkit/issues
|
|
8
|
+
Requires-Python: >=3.10
|
|
9
|
+
Description-Content-Type: text/markdown
|
|
10
|
+
Requires-Dist: pandas
|
|
11
|
+
Requires-Dist: numpy
|
|
12
|
+
Requires-Dist: matplotlib
|
|
13
|
+
Requires-Dist: seaborn
|
|
14
|
+
Provides-Extra: ml
|
|
15
|
+
Requires-Dist: scikit-learn; extra == "ml"
|
|
16
|
+
|
|
17
|
+
# ml-gearbox
|
|
18
|
+
|
|
19
|
+
A reusable set of pandas-based functions for two jobs every ML project needs before modeling:
|
|
20
|
+
|
|
21
|
+
- **`audit`** — is this dataset trustworthy? (missingness, duplicates, schema, validity)
|
|
22
|
+
- **`eda`** — what does this dataset actually look like, and how does it relate to the target?
|
|
23
|
+
|
|
24
|
+
This README documents every function currently in the toolkit: what it checks, what it expects from you, what it returns, and — most importantly — **what decision to make once you see the result.** It's meant to be read like a reference manual, not front-to-back.
|
|
25
|
+
|
|
26
|
+
## Install
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
pip install -e .
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
(from the project root, where `pyproject.toml` lives). This installs the package in "editable" mode, so changes to the source are picked up immediately without reinstalling.
|
|
33
|
+
|
|
34
|
+
## Quickstart
|
|
35
|
+
|
|
36
|
+
```python
|
|
37
|
+
from ml_gearbox.audit import run_data_audit
|
|
38
|
+
from ml_gearbox.eda import iqr_outlier_summary, plot_correlation_heatmap, correlation_matrix
|
|
39
|
+
|
|
40
|
+
# Step 1: is the data trustworthy?
|
|
41
|
+
audit = run_data_audit(df, target_col="default", id_col="customer_id")
|
|
42
|
+
print(audit["missingness"])
|
|
43
|
+
print(audit["duplicates"])
|
|
44
|
+
|
|
45
|
+
# Step 2: once it's clean, explore it
|
|
46
|
+
outliers = iqr_outlier_summary(df, columns=["age", "income"])
|
|
47
|
+
corr = correlation_matrix(df, columns=["age", "income", "tenure"])
|
|
48
|
+
plot_correlation_heatmap(corr)
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
## How to think about the two modules
|
|
52
|
+
|
|
53
|
+
Run **`audit` first, always**. It answers "can I trust this data at all?" Nothing in `eda` is meaningful if your dataset has silent duplicate IDs, 40% hidden missingness, or a target column that's actually a string when you expect numeric.
|
|
54
|
+
|
|
55
|
+
Once audit comes back clean (or you've fixed what it found), move to **`eda`** to understand *shape*: distributions, outliers, correlations, and how features relate to your target.
|
|
56
|
+
|
|
57
|
+
### Suggested order of operations
|
|
58
|
+
|
|
59
|
+
1. `schema_report` — get the lay of the land (dtypes, missingness, uniqueness, in one table)
|
|
60
|
+
2. `id_report` + `duplicate_report` — is each row a genuine, unique observation?
|
|
61
|
+
3. `missingness_report` + `hidden_missing_report` — where are the gaps, including disguised ones?
|
|
62
|
+
4. `category_normalization_check` — are your categories cleaner than they look?
|
|
63
|
+
5. `infinite_value_report` — any `inf` values that will silently break downstream math?
|
|
64
|
+
6. `cardinality_report` + `constant_feature_report` — anything uninformative or unmanageably high-cardinality?
|
|
65
|
+
7. `schema_validation` / `range_validation` / `cross_field_validation` — once you know the rules this dataset should follow, check it actually follows them
|
|
66
|
+
8. Now move to `eda.py` — distributions, outliers, correlations, target rates, and their plots
|
|
67
|
+
|
|
68
|
+
Or just call `run_data_audit(...)` to get steps 1–7 in one dictionary.
|
|
69
|
+
|
|
70
|
+
---
|
|
71
|
+
|
|
72
|
+
## `audit.py` reference
|
|
73
|
+
|
|
74
|
+
### Target, ID, and duplicates
|
|
75
|
+
|
|
76
|
+
#### `target_report(df, target_col) -> DataFrame`
|
|
77
|
+
- **Checks:** the class balance of your target variable.
|
|
78
|
+
- **Expects:** `target_col` is a column name in `df`. Works for any number of classes, including a numeric binary target.
|
|
79
|
+
- **Returns:** one row per class, with `count` and `percentage`.
|
|
80
|
+
- **Act on it:** if one class is under ~5–10%, you're dealing with class imbalance — plan for it now (stratified splits, class weights, resampling, or a metric other than accuracy) rather than after your first model mysteriously predicts one class every time.
|
|
81
|
+
|
|
82
|
+
#### `id_report(df, id_col) -> dict`
|
|
83
|
+
- **Checks:** whether your identifier column is actually a valid identifier.
|
|
84
|
+
- **Expects:** `id_col` is the column meant to uniquely identify each row (a customer ID, transaction ID, etc.).
|
|
85
|
+
- **Returns:** `total_rows`, `unique_ids`, `missing_ids`, `duplicate_ids`.
|
|
86
|
+
- **Act on it:** `missing_ids > 0` or `duplicate_ids > 0` means something is wrong upstream — a broken join, a bad extract, or two systems using the same ID space differently. Investigate the source before doing anything else; don't just drop rows and move on.
|
|
87
|
+
|
|
88
|
+
#### `duplicate_report(df, id_col=None) -> dict`
|
|
89
|
+
- **Checks:** duplicate records at two levels. `full_row_duplicates` counts rows that are identical across *every* column (safe to drop — almost always a copy-paste or re-import artifact). If you pass `id_col`, it additionally reports `duplicate_id_rows` (rows sharing an ID) and, critically, `conflicting_id_rows` — rows that share the same ID but have **different** values elsewhere.
|
|
90
|
+
- **Expects:** `id_col` is optional but strongly recommended if your dataset has one; without it you only get the full-row check, which will show 0 on almost any dataset that has an ID column (since the ID makes every row technically unique).
|
|
91
|
+
- **Returns:** a dict, e.g. `{"full_row_duplicates": 1, "duplicate_id_rows": 2, "conflicting_id_rows": 2}`.
|
|
92
|
+
- **Act on it:** `full_row_duplicates` → safe to `df.drop_duplicates()`. `conflicting_id_rows` is the dangerous one — it means the same entity has two different versions of the truth on record (e.g. two different incomes for the same customer ID). Don't silently pick one; find out which row is current/correct, usually via a timestamp column if one exists.
|
|
93
|
+
|
|
94
|
+
### Missingness
|
|
95
|
+
|
|
96
|
+
#### `missingness_report(df) -> DataFrame`
|
|
97
|
+
- **Checks:** standard missing values (`NaN`/`None`) per column.
|
|
98
|
+
- **Expects:** nothing special — works on the whole dataframe.
|
|
99
|
+
- **Returns:** `feature`, `missing_count`, `percentage`, only for columns with at least one missing value, sorted worst-first. Empty DataFrame if nothing's missing.
|
|
100
|
+
- **Act on it:** under ~5% missing → usually safe to impute (median/mode) or drop rows. 5–40% → impute carefully, consider a "was_missing" flag as its own feature since the *fact* of missingness can be predictive. Above ~40–50% → seriously consider dropping the column; imputation is mostly guessing at that point.
|
|
101
|
+
|
|
102
|
+
#### `hidden_missing_report(df) -> DataFrame`
|
|
103
|
+
- **Checks:** missingness that doesn't show up in `.isna()` because it's disguised as an empty or whitespace-only string (`""`, `" "`).
|
|
104
|
+
- **Expects:** runs on object/category columns only.
|
|
105
|
+
- **Returns:** `feature`, `hidden_missing_count`, `percentage`. Empty DataFrame if none found.
|
|
106
|
+
- **Act on it:** any row this flags should be treated as missing, not as a valid (blank) category. Convert them to real `NaN` (`df[col] = df[col].replace(r'^\s*$', pd.NA, regex=True)`) before running `missingness_report` again, or your missingness numbers upstream are undercounting.
|
|
107
|
+
|
|
108
|
+
#### `category_normalization_check(df, columns) -> DataFrame`
|
|
109
|
+
- **Checks:** whether categories that look different are actually the same thing with different formatting — `"USA"`, `"usa"`, `" USA "` all collapsing into one category once stripped and lowercased.
|
|
110
|
+
- **Expects:** a list of categorical column names to check.
|
|
111
|
+
- **Returns:** `feature`, `original_unique`, `normalized_unique` — only for columns where normalizing *reduces* the category count. Empty DataFrame if formatting is already clean.
|
|
112
|
+
- **Act on it:** if a column shows up here, your `cardinality_report` and any category-based model feature (one-hot encoding, target encoding) is currently splitting one real-world category into several. Normalize the column (`.str.strip().str.lower()`, or a proper mapping) before encoding, not after.
|
|
113
|
+
|
|
114
|
+
#### `infinite_value_report(df) -> DataFrame`
|
|
115
|
+
- **Checks:** `inf` / `-inf` values in numeric columns — usually the result of a division by zero somewhere upstream (e.g. a ratio feature like `income / years_employed` where `years_employed` was 0).
|
|
116
|
+
- **Expects:** nothing special.
|
|
117
|
+
- **Returns:** `feature`, `infinite_count`. Empty DataFrame if none found.
|
|
118
|
+
- **Act on it:** treat these as missing (`df.replace([np.inf, -np.inf], np.nan)`) or fix the calculation that produced them. **Do this before running any quantile-based function** (see the Known Quirks section below) — `inf` values will crash `numeric_target_rate_by_quantile` and `pd.qcut`-based binning in general.
|
|
119
|
+
|
|
120
|
+
### Schema, cardinality, constants
|
|
121
|
+
|
|
122
|
+
#### `schema_report(df) -> DataFrame`
|
|
123
|
+
- **Checks:** a one-table overview of every column: dtype, missing count/%, and unique count/%.
|
|
124
|
+
- **Expects:** nothing special. This is usually your first call on a new dataset.
|
|
125
|
+
- **Returns:** one row per column, sorted by missingness (worst first).
|
|
126
|
+
- **Act on it:** use this as your map. Anything with `dtype == object` that you expected to be numeric, or `unique_pct` near 100% on a column that isn't an ID, deserves a closer look before you go further.
|
|
127
|
+
|
|
128
|
+
#### `cardinality_report(df) -> DataFrame`
|
|
129
|
+
- **Checks:** how many unique values each categorical (object/category) column has.
|
|
130
|
+
- **Expects:** nothing special.
|
|
131
|
+
- **Returns:** `feature`, `n_unique`, sorted highest first.
|
|
132
|
+
- **Act on it:** columns with very high cardinality (hundreds/thousands of unique values — free-text names, raw IDs accidentally typed as category) will blow up one-hot encoding. Use target/frequency encoding, group rare categories into "other", or drop the column if it's really just an identifier.
|
|
133
|
+
|
|
134
|
+
#### `constant_feature_report(df) -> DataFrame`
|
|
135
|
+
- **Checks:** how dominant the single most common value is in each column (works on *any* column, not just categorical).
|
|
136
|
+
- **Expects:** nothing special.
|
|
137
|
+
- **Returns:** `feature`, `dominant_percentage`, `is_strictly_constant` (True only if one value makes up 100%), sorted highest first.
|
|
138
|
+
- **Act on it:** `is_strictly_constant == True` → the column carries zero information for modeling; drop it. Dominant but not strictly constant (e.g. 97%) → the column will likely have very low predictive power and near-zero variance; consider dropping it too, or check whether the rare values are actually meaningful outliers worth keeping.
|
|
139
|
+
|
|
140
|
+
#### `numeric_profile(df) -> DataFrame`
|
|
141
|
+
- **Checks:** an enhanced `.describe()` for numeric columns — adds median and skew on top of the standard mean/std/min/max/quartiles.
|
|
142
|
+
- **Expects:** nothing special. Returns an empty DataFrame if there are no numeric columns.
|
|
143
|
+
- **Returns:** one row per numeric column, sorted by skew (most skewed first).
|
|
144
|
+
- **Act on it:** high absolute skew (roughly beyond ±1) means a log or power transform will likely help linear models and distance-based methods; tree-based models (random forest, gradient boosting) don't care about skew and can be left alone. A big gap between `mean` and `median` is another sign of skew or outlier influence.
|
|
145
|
+
|
|
146
|
+
### Validation (these need a spec from you)
|
|
147
|
+
|
|
148
|
+
These three are different from everything above: they don't infer what's "wrong" automatically — **you** tell them the rules for this specific dataset, and they check the data against those rules. This can't be fully generic (a negative age is invalid; a negative account balance might be completely normal), but the mechanism for checking is reusable across any dataset.
|
|
149
|
+
|
|
150
|
+
#### `schema_validation(df, expected_schema) -> DataFrame`
|
|
151
|
+
- **Checks:** does the dataframe match a dtype contract you define?
|
|
152
|
+
- **Expects:** `expected_schema` — a dict of `{column_name: expected_dtype}`, e.g. `{"age": "int64", "income": "float64", "country": "object"}`.
|
|
153
|
+
- **Returns:** one row per column (expected or actual), with `status` of `ok`, `dtype_mismatch`, `missing_column` (you expected it, it's not there), or `unexpected_column` (it's there, you didn't account for it).
|
|
154
|
+
- **Act on it:** `missing_column` usually means an upstream extract or join dropped a field — treat as a pipeline bug. `dtype_mismatch` (e.g. you expected `float64` but got `object`) is a classic sign that a stray non-numeric value (a typo, a placeholder string like `"N/A"`) is silently corrupting the whole column's type — find and fix that value before modeling. `unexpected_column` isn't necessarily bad, just update your schema once you've confirmed it's intentional.
|
|
155
|
+
|
|
156
|
+
#### `range_validation(df, rules) -> DataFrame`
|
|
157
|
+
- **Checks:** numeric bounds and/or category whitelists, in one pass.
|
|
158
|
+
- **Expects:** `rules` — a dict of `{column_name: constraint_dict}`, where each constraint dict can have `"min"`, `"max"`, and/or `"allowed"` (a list of permitted values). Example: `{"age": {"min": 0, "max": 120}, "status": {"allowed": ["active", "closed", "pending"]}}`.
|
|
159
|
+
- **Returns:** one row per constraint you defined (a column with both `min` and `max` produces two rows), with `violation_count` and `violation_pct`.
|
|
160
|
+
- **Act on it:** any violation count above 0 needs a judgment call: is this bad data (a negative age → almost certainly an error, fix or drop it) or a rule that was wrong (maybe -1 is a legitimate "unknown" sentinel value in this dataset → adjust the rule, not the data). Don't silently clip or drop without checking which case you're in.
|
|
161
|
+
|
|
162
|
+
#### `cross_field_validation(df, rules) -> DataFrame`
|
|
163
|
+
- **Checks:** logical relationships *between* two columns in the same row.
|
|
164
|
+
- **Expects:** `rules` — a list of `(col_a, operator, col_b)` tuples. Supported operators: `>=`, `>`, `<=`, `<`, `==`, `!=`. Example: `[("end_date", ">=", "start_date"), ("net_income", "<=", "gross_income")]`. Rows where either column is missing are skipped, not flagged.
|
|
165
|
+
- **Returns:** one row per rule, with `violation_count` and `violation_pct`.
|
|
166
|
+
- **Act on it:** this is one of the highest-signal checks in the toolkit — a violation here usually means a genuine data entry or pipeline error (an end date before a start date isn't a modeling nuisance, it's a broken record). Pull the violating rows out and inspect them individually rather than aggregating past them.
|
|
167
|
+
|
|
168
|
+
### Master wrapper
|
|
169
|
+
|
|
170
|
+
#### `run_data_audit(df, target_col, id_col, expected_schema=None, range_rules=None, cross_field_rules=None) -> dict`
|
|
171
|
+
- **Checks:** everything above, in one call.
|
|
172
|
+
- **Expects:** `target_col` and `id_col` are required. `expected_schema`, `range_rules`, `cross_field_rules` are optional — pass them once you know what "valid" means for this dataset, and their reports get added to the output automatically.
|
|
173
|
+
- **Returns:** a dict keyed by report name (`"schema"`, `"target"`, `"id"`, `"duplicates"`, `"missingness"`, `"hidden_missingness"`, `"category_normalization"`, `"infinite_values"`, `"cardinality"`, `"constant_features"`, `"numeric_profile"`, plus the three validation reports if you supplied rules for them).
|
|
174
|
+
- **Act on it:** treat this as your standard "first thing I run on any new dataset" call. Loop through the dict and eyeball each report before writing a single line of feature engineering.
|
|
175
|
+
|
|
176
|
+
---
|
|
177
|
+
|
|
178
|
+
## `eda.py` reference
|
|
179
|
+
|
|
180
|
+
### Distributions
|
|
181
|
+
|
|
182
|
+
#### `target_distribution(df, target) -> DataFrame`
|
|
183
|
+
- **Checks:** the same thing as `target_report` in `audit.py` — count and percentage per class of the target.
|
|
184
|
+
- **Expects:** `target` is a column name.
|
|
185
|
+
- **Returns:** one row per class.
|
|
186
|
+
- **Act on it:** same as `target_report` — use it to catch class imbalance early. (These two functions currently duplicate each other; a future cleanup will likely have `eda` call `audit`'s version instead of recomputing it.)
|
|
187
|
+
|
|
188
|
+
#### `zero_percentage(df, columns) -> DataFrame`
|
|
189
|
+
- **Checks:** what fraction of each numeric column is exactly `0` (as opposed to missing).
|
|
190
|
+
- **Expects:** a list of numeric column names.
|
|
191
|
+
- **Returns:** `feature`, `zero_pct`, only for columns with at least one zero, sorted highest first.
|
|
192
|
+
- **Act on it:** a high zero percentage isn't automatically a problem — sometimes zero is a genuine value (zero purchases, zero late payments). But if it's unexpectedly high for a column like "income" or "age", it may actually be a placeholder for missing data that never got converted to `NaN`. Check the data dictionary or source system before treating zeros at face value.
|
|
193
|
+
|
|
194
|
+
### Outliers
|
|
195
|
+
|
|
196
|
+
#### `iqr_outlier_summary(df, columns) -> DataFrame`
|
|
197
|
+
- **Checks:** outliers per numeric column using the standard IQR rule (outside `[Q1 - 1.5×IQR, Q3 + 1.5×IQR]`).
|
|
198
|
+
- **Expects:** a list of numeric column names.
|
|
199
|
+
- **Returns:** `feature`, `q1`, `q3`, `iqr`, `lower_bound`, `upper_bound`, `outlier_count`, `outlier_pct`, sorted worst first.
|
|
200
|
+
- **Act on it:** a handful of outliers (under ~1-2%) is normal — look at a few individually to confirm they're real, not data entry errors, and decide whether to cap/winsorize, transform (log), or leave them (tree-based models are largely outlier-robust; linear/distance-based models are not). A very high outlier percentage usually means the IQR rule doesn't fit this column's distribution (e.g. it's naturally right-skewed, like income) rather than that 20% of your data is broken — pair this with `numeric_profile`'s skew column before reacting.
|
|
201
|
+
|
|
202
|
+
### Target relationships
|
|
203
|
+
|
|
204
|
+
#### `categorical_target_rate(df, feature, target) -> DataFrame`
|
|
205
|
+
- **Checks:** the target rate (e.g. default rate, churn rate) within each category of a feature.
|
|
206
|
+
- **Expects:** `target` should be numeric 0/1 (the function calls `.mean()` on it, which needs a numeric target).
|
|
207
|
+
- **Returns:** `feature` (category), `count`, `positives`, `target_rate`, `target_rate_pct`, sorted highest rate first.
|
|
208
|
+
- **Act on it:** big spread in target rate across categories → this feature likely has real predictive power, keep it. Also check `count` per category — a category with a dramatic rate but only 3 rows is noise, not signal; consider grouping small categories together.
|
|
209
|
+
|
|
210
|
+
#### `numeric_target_rate_by_quantile(df, feature, target, bins=10) -> DataFrame`
|
|
211
|
+
- **Checks:** the same idea as above, but for a continuous feature — splits it into quantile bins and shows the target rate per bin.
|
|
212
|
+
- **Expects:** `target` numeric 0/1. **Important:** the feature column must not contain `inf` values (see Known Quirks below) — run `infinite_value_report` first and clean those up.
|
|
213
|
+
- **Returns:** `bin` (the quantile range), `count`, `positives`, `target_rate`, `target_rate_pct`.
|
|
214
|
+
- **Act on it:** a target rate that rises or falls monotonically across bins → strong signal, this feature (or a transformed/binned version of it) will likely help a model. A non-monotonic, jagged pattern (e.g. high-low-high-low) → a linear model will miss this relationship even though it's real; a tree-based model or explicit binning will capture it better.
|
|
215
|
+
|
|
216
|
+
### Correlations
|
|
217
|
+
|
|
218
|
+
#### `correlation_matrix(df, columns, method="pearson") -> DataFrame`
|
|
219
|
+
- **Checks:** pairwise correlation between numeric columns.
|
|
220
|
+
- **Expects:** a list of numeric columns; `method` can be `"pearson"` (linear), `"spearman"` (monotonic/rank-based), or `"kendall"`.
|
|
221
|
+
- **Returns:** a square correlation matrix.
|
|
222
|
+
- **Act on it:** this is usually consumed by `high_correlation_pairs` and `plot_correlation_heatmap` rather than read raw — see below.
|
|
223
|
+
|
|
224
|
+
#### `high_correlation_pairs(corr_matrix, threshold=0.85) -> DataFrame`
|
|
225
|
+
- **Checks:** pulls out just the feature pairs from a correlation matrix that exceed a given absolute correlation.
|
|
226
|
+
- **Expects:** the output of `correlation_matrix(...)`.
|
|
227
|
+
- **Returns:** `feature_1`, `feature_2`, `correlation`, sorted by absolute strength. Empty DataFrame if nothing crosses the threshold.
|
|
228
|
+
- **Act on it:** pairs above ~0.85–0.9 are close to redundant. For linear/logistic regression, this kind of multicollinearity inflates coefficient variance and makes them unstable — drop one feature from each pair, or combine them (e.g. a ratio or PCA component). Tree-based models tolerate this better, but you're still paying for two features' worth of noise for one feature's worth of signal.
|
|
229
|
+
|
|
230
|
+
### Plots
|
|
231
|
+
|
|
232
|
+
All plotting functions call `plt.show()` and don't return a figure — they're meant for interactive/notebook use. Each one mirrors a data function above so you get the numbers and the picture together.
|
|
233
|
+
|
|
234
|
+
#### `plot_numeric_distribution(df, column)`
|
|
235
|
+
- **Shows:** a histogram (with KDE) and a boxplot side by side for one numeric column.
|
|
236
|
+
- **Use it after:** `numeric_profile` flags a column with high skew, or `iqr_outlier_summary` flags one with a high outlier percentage — this lets you see the shape, not just the number.
|
|
237
|
+
|
|
238
|
+
#### `plot_categorical_distribution(df, column, top_n=15)`
|
|
239
|
+
- **Shows:** a horizontal bar chart of category frequency for a categorical column (capped at the top N categories, default 15, so it stays readable on high-cardinality columns).
|
|
240
|
+
- **Use it after:** `cardinality_report` flags a column — visually confirm whether it's a few dominant categories with a long thin tail (a good candidate for "group rare into Other"), or genuinely spread out.
|
|
241
|
+
|
|
242
|
+
#### `plot_numeric_by_target(df, feature, target)`
|
|
243
|
+
- **Shows:** overlapping density curves of a numeric feature, split by target class.
|
|
244
|
+
- **Use it after:** you want to eyeball whether a feature separates the two classes at all, before formally checking with `numeric_target_rate_by_quantile`. Curves that barely overlap → strong feature. Curves that sit almost on top of each other → weak feature.
|
|
245
|
+
|
|
246
|
+
#### `plot_categorical_target_rate(df, feature, target)`
|
|
247
|
+
- **Shows:** the same information as `categorical_target_rate`, as a horizontal bar chart.
|
|
248
|
+
- **Use it after:** calling `categorical_target_rate` — bars make it much faster to spot which categories are driving risk/behavior than scanning a table.
|
|
249
|
+
|
|
250
|
+
#### `plot_numeric_target_rate(df, feature, target, bins=10)`
|
|
251
|
+
- **Shows:** the same information as `numeric_target_rate_by_quantile`, as a line plot across bins.
|
|
252
|
+
- **Use it after:** calling `numeric_target_rate_by_quantile` — the line makes monotonic vs. jagged relationships immediately visible (see the "Act on it" note above on why that distinction matters for model choice). Same `inf`-value caveat applies.
|
|
253
|
+
|
|
254
|
+
#### `plot_correlation_heatmap(corr_matrix)`
|
|
255
|
+
- **Shows:** a heatmap of a correlation matrix, annotated with the actual coefficient values.
|
|
256
|
+
- **Use it after:** calling `correlation_matrix` — faster than scanning `high_correlation_pairs` when you want the full picture of how everything relates to everything, not just the pairs above a threshold.
|
|
257
|
+
|
|
258
|
+
---
|
|
259
|
+
|
|
260
|
+
## Known quirks
|
|
261
|
+
|
|
262
|
+
- **`numeric_target_rate_by_quantile` (and its plot) will crash on `inf` values.** `pd.qcut` can't build bin edges around an infinite value. This is a real interaction between the two modules: `infinite_value_report` exists specifically to catch this kind of thing, so make it a habit to run it — and clean up whatever it finds — before reaching for any quantile-based EDA function.
|
|
263
|
+
- **`target_distribution` (eda.py) and `target_report` (audit.py) do the same calculation.** Harmless for now, but if you extend one, extend both, or better, consolidate them the next time you touch this code.
|
|
264
|
+
- **`schema_validation`'s dtype comparison is exact.** Depending on your pandas version, a plain string column may report as dtype `object` or as a newer string-specific dtype. If you get unexpected `dtype_mismatch` rows on columns that look fine, check what `df[col].dtype` actually prints in your environment and adjust `expected_schema` to match — this isn't a bug, it's the check doing its job of catching a dtype you didn't expect.
|
|
265
|
+
|
|
266
|
+
## Extending this toolkit
|
|
267
|
+
|
|
268
|
+
When adding a new function, keep the two existing conventions:
|
|
269
|
+
1. **Data functions return a DataFrame (or dict)**, never a plot. Keep plotting in separate `plot_*` functions that consume the data function's output where possible — that's what makes each check independently testable and reusable outside of a notebook.
|
|
270
|
+
2. **Document it here in the same format**: what it checks, what it expects, what it returns, and what decision to make from the result. A check nobody knows how to act on isn't pulling its weight in a toolkit like this one.
|
|
@@ -0,0 +1,254 @@
|
|
|
1
|
+
# ml-gearbox
|
|
2
|
+
|
|
3
|
+
A reusable set of pandas-based functions for two jobs every ML project needs before modeling:
|
|
4
|
+
|
|
5
|
+
- **`audit`** — is this dataset trustworthy? (missingness, duplicates, schema, validity)
|
|
6
|
+
- **`eda`** — what does this dataset actually look like, and how does it relate to the target?
|
|
7
|
+
|
|
8
|
+
This README documents every function currently in the toolkit: what it checks, what it expects from you, what it returns, and — most importantly — **what decision to make once you see the result.** It's meant to be read like a reference manual, not front-to-back.
|
|
9
|
+
|
|
10
|
+
## Install
|
|
11
|
+
|
|
12
|
+
```bash
|
|
13
|
+
pip install -e .
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
(from the project root, where `pyproject.toml` lives). This installs the package in "editable" mode, so changes to the source are picked up immediately without reinstalling.
|
|
17
|
+
|
|
18
|
+
## Quickstart
|
|
19
|
+
|
|
20
|
+
```python
|
|
21
|
+
from ml_gearbox.audit import run_data_audit
|
|
22
|
+
from ml_gearbox.eda import iqr_outlier_summary, plot_correlation_heatmap, correlation_matrix
|
|
23
|
+
|
|
24
|
+
# Step 1: is the data trustworthy?
|
|
25
|
+
audit = run_data_audit(df, target_col="default", id_col="customer_id")
|
|
26
|
+
print(audit["missingness"])
|
|
27
|
+
print(audit["duplicates"])
|
|
28
|
+
|
|
29
|
+
# Step 2: once it's clean, explore it
|
|
30
|
+
outliers = iqr_outlier_summary(df, columns=["age", "income"])
|
|
31
|
+
corr = correlation_matrix(df, columns=["age", "income", "tenure"])
|
|
32
|
+
plot_correlation_heatmap(corr)
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
## How to think about the two modules
|
|
36
|
+
|
|
37
|
+
Run **`audit` first, always**. It answers "can I trust this data at all?" Nothing in `eda` is meaningful if your dataset has silent duplicate IDs, 40% hidden missingness, or a target column that's actually a string when you expect numeric.
|
|
38
|
+
|
|
39
|
+
Once audit comes back clean (or you've fixed what it found), move to **`eda`** to understand *shape*: distributions, outliers, correlations, and how features relate to your target.
|
|
40
|
+
|
|
41
|
+
### Suggested order of operations
|
|
42
|
+
|
|
43
|
+
1. `schema_report` — get the lay of the land (dtypes, missingness, uniqueness, in one table)
|
|
44
|
+
2. `id_report` + `duplicate_report` — is each row a genuine, unique observation?
|
|
45
|
+
3. `missingness_report` + `hidden_missing_report` — where are the gaps, including disguised ones?
|
|
46
|
+
4. `category_normalization_check` — are your categories cleaner than they look?
|
|
47
|
+
5. `infinite_value_report` — any `inf` values that will silently break downstream math?
|
|
48
|
+
6. `cardinality_report` + `constant_feature_report` — anything uninformative or unmanageably high-cardinality?
|
|
49
|
+
7. `schema_validation` / `range_validation` / `cross_field_validation` — once you know the rules this dataset should follow, check it actually follows them
|
|
50
|
+
8. Now move to `eda.py` — distributions, outliers, correlations, target rates, and their plots
|
|
51
|
+
|
|
52
|
+
Or just call `run_data_audit(...)` to get steps 1–7 in one dictionary.
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## `audit.py` reference
|
|
57
|
+
|
|
58
|
+
### Target, ID, and duplicates
|
|
59
|
+
|
|
60
|
+
#### `target_report(df, target_col) -> DataFrame`
|
|
61
|
+
- **Checks:** the class balance of your target variable.
|
|
62
|
+
- **Expects:** `target_col` is a column name in `df`. Works for any number of classes, including a numeric binary target.
|
|
63
|
+
- **Returns:** one row per class, with `count` and `percentage`.
|
|
64
|
+
- **Act on it:** if one class is under ~5–10%, you're dealing with class imbalance — plan for it now (stratified splits, class weights, resampling, or a metric other than accuracy) rather than after your first model mysteriously predicts one class every time.
|
|
65
|
+
|
|
66
|
+
#### `id_report(df, id_col) -> dict`
|
|
67
|
+
- **Checks:** whether your identifier column is actually a valid identifier.
|
|
68
|
+
- **Expects:** `id_col` is the column meant to uniquely identify each row (a customer ID, transaction ID, etc.).
|
|
69
|
+
- **Returns:** `total_rows`, `unique_ids`, `missing_ids`, `duplicate_ids`.
|
|
70
|
+
- **Act on it:** `missing_ids > 0` or `duplicate_ids > 0` means something is wrong upstream — a broken join, a bad extract, or two systems using the same ID space differently. Investigate the source before doing anything else; don't just drop rows and move on.
|
|
71
|
+
|
|
72
|
+
#### `duplicate_report(df, id_col=None) -> dict`
|
|
73
|
+
- **Checks:** duplicate records at two levels. `full_row_duplicates` counts rows that are identical across *every* column (safe to drop — almost always a copy-paste or re-import artifact). If you pass `id_col`, it additionally reports `duplicate_id_rows` (rows sharing an ID) and, critically, `conflicting_id_rows` — rows that share the same ID but have **different** values elsewhere.
|
|
74
|
+
- **Expects:** `id_col` is optional but strongly recommended if your dataset has one; without it you only get the full-row check, which will show 0 on almost any dataset that has an ID column (since the ID makes every row technically unique).
|
|
75
|
+
- **Returns:** a dict, e.g. `{"full_row_duplicates": 1, "duplicate_id_rows": 2, "conflicting_id_rows": 2}`.
|
|
76
|
+
- **Act on it:** `full_row_duplicates` → safe to `df.drop_duplicates()`. `conflicting_id_rows` is the dangerous one — it means the same entity has two different versions of the truth on record (e.g. two different incomes for the same customer ID). Don't silently pick one; find out which row is current/correct, usually via a timestamp column if one exists.
|
|
77
|
+
|
|
78
|
+
### Missingness
|
|
79
|
+
|
|
80
|
+
#### `missingness_report(df) -> DataFrame`
|
|
81
|
+
- **Checks:** standard missing values (`NaN`/`None`) per column.
|
|
82
|
+
- **Expects:** nothing special — works on the whole dataframe.
|
|
83
|
+
- **Returns:** `feature`, `missing_count`, `percentage`, only for columns with at least one missing value, sorted worst-first. Empty DataFrame if nothing's missing.
|
|
84
|
+
- **Act on it:** under ~5% missing → usually safe to impute (median/mode) or drop rows. 5–40% → impute carefully, consider a "was_missing" flag as its own feature since the *fact* of missingness can be predictive. Above ~40–50% → seriously consider dropping the column; imputation is mostly guessing at that point.
|
|
85
|
+
|
|
86
|
+
#### `hidden_missing_report(df) -> DataFrame`
|
|
87
|
+
- **Checks:** missingness that doesn't show up in `.isna()` because it's disguised as an empty or whitespace-only string (`""`, `" "`).
|
|
88
|
+
- **Expects:** runs on object/category columns only.
|
|
89
|
+
- **Returns:** `feature`, `hidden_missing_count`, `percentage`. Empty DataFrame if none found.
|
|
90
|
+
- **Act on it:** any row this flags should be treated as missing, not as a valid (blank) category. Convert them to real `NaN` (`df[col] = df[col].replace(r'^\s*$', pd.NA, regex=True)`) before running `missingness_report` again, or your missingness numbers upstream are undercounting.
|
|
91
|
+
|
|
92
|
+
#### `category_normalization_check(df, columns) -> DataFrame`
|
|
93
|
+
- **Checks:** whether categories that look different are actually the same thing with different formatting — `"USA"`, `"usa"`, `" USA "` all collapsing into one category once stripped and lowercased.
|
|
94
|
+
- **Expects:** a list of categorical column names to check.
|
|
95
|
+
- **Returns:** `feature`, `original_unique`, `normalized_unique` — only for columns where normalizing *reduces* the category count. Empty DataFrame if formatting is already clean.
|
|
96
|
+
- **Act on it:** if a column shows up here, your `cardinality_report` and any category-based model feature (one-hot encoding, target encoding) is currently splitting one real-world category into several. Normalize the column (`.str.strip().str.lower()`, or a proper mapping) before encoding, not after.
|
|
97
|
+
|
|
98
|
+
#### `infinite_value_report(df) -> DataFrame`
|
|
99
|
+
- **Checks:** `inf` / `-inf` values in numeric columns — usually the result of a division by zero somewhere upstream (e.g. a ratio feature like `income / years_employed` where `years_employed` was 0).
|
|
100
|
+
- **Expects:** nothing special.
|
|
101
|
+
- **Returns:** `feature`, `infinite_count`. Empty DataFrame if none found.
|
|
102
|
+
- **Act on it:** treat these as missing (`df.replace([np.inf, -np.inf], np.nan)`) or fix the calculation that produced them. **Do this before running any quantile-based function** (see the Known Quirks section below) — `inf` values will crash `numeric_target_rate_by_quantile` and `pd.qcut`-based binning in general.
|
|
103
|
+
|
|
104
|
+
### Schema, cardinality, constants
|
|
105
|
+
|
|
106
|
+
#### `schema_report(df) -> DataFrame`
|
|
107
|
+
- **Checks:** a one-table overview of every column: dtype, missing count/%, and unique count/%.
|
|
108
|
+
- **Expects:** nothing special. This is usually your first call on a new dataset.
|
|
109
|
+
- **Returns:** one row per column, sorted by missingness (worst first).
|
|
110
|
+
- **Act on it:** use this as your map. Anything with `dtype == object` that you expected to be numeric, or `unique_pct` near 100% on a column that isn't an ID, deserves a closer look before you go further.
|
|
111
|
+
|
|
112
|
+
#### `cardinality_report(df) -> DataFrame`
|
|
113
|
+
- **Checks:** how many unique values each categorical (object/category) column has.
|
|
114
|
+
- **Expects:** nothing special.
|
|
115
|
+
- **Returns:** `feature`, `n_unique`, sorted highest first.
|
|
116
|
+
- **Act on it:** columns with very high cardinality (hundreds/thousands of unique values — free-text names, raw IDs accidentally typed as category) will blow up one-hot encoding. Use target/frequency encoding, group rare categories into "other", or drop the column if it's really just an identifier.
|
|
117
|
+
|
|
118
|
+
#### `constant_feature_report(df) -> DataFrame`
|
|
119
|
+
- **Checks:** how dominant the single most common value is in each column (works on *any* column, not just categorical).
|
|
120
|
+
- **Expects:** nothing special.
|
|
121
|
+
- **Returns:** `feature`, `dominant_percentage`, `is_strictly_constant` (True only if one value makes up 100%), sorted highest first.
|
|
122
|
+
- **Act on it:** `is_strictly_constant == True` → the column carries zero information for modeling; drop it. Dominant but not strictly constant (e.g. 97%) → the column will likely have very low predictive power and near-zero variance; consider dropping it too, or check whether the rare values are actually meaningful outliers worth keeping.
|
|
123
|
+
|
|
124
|
+
#### `numeric_profile(df) -> DataFrame`
|
|
125
|
+
- **Checks:** an enhanced `.describe()` for numeric columns — adds median and skew on top of the standard mean/std/min/max/quartiles.
|
|
126
|
+
- **Expects:** nothing special. Returns an empty DataFrame if there are no numeric columns.
|
|
127
|
+
- **Returns:** one row per numeric column, sorted by skew (most skewed first).
|
|
128
|
+
- **Act on it:** high absolute skew (roughly beyond ±1) means a log or power transform will likely help linear models and distance-based methods; tree-based models (random forest, gradient boosting) don't care about skew and can be left alone. A big gap between `mean` and `median` is another sign of skew or outlier influence.
|
|
129
|
+
|
|
130
|
+
### Validation (these need a spec from you)
|
|
131
|
+
|
|
132
|
+
These three are different from everything above: they don't infer what's "wrong" automatically — **you** tell them the rules for this specific dataset, and they check the data against those rules. This can't be fully generic (a negative age is invalid; a negative account balance might be completely normal), but the mechanism for checking is reusable across any dataset.
|
|
133
|
+
|
|
134
|
+
#### `schema_validation(df, expected_schema) -> DataFrame`
|
|
135
|
+
- **Checks:** does the dataframe match a dtype contract you define?
|
|
136
|
+
- **Expects:** `expected_schema` — a dict of `{column_name: expected_dtype}`, e.g. `{"age": "int64", "income": "float64", "country": "object"}`.
|
|
137
|
+
- **Returns:** one row per column (expected or actual), with `status` of `ok`, `dtype_mismatch`, `missing_column` (you expected it, it's not there), or `unexpected_column` (it's there, you didn't account for it).
|
|
138
|
+
- **Act on it:** `missing_column` usually means an upstream extract or join dropped a field — treat as a pipeline bug. `dtype_mismatch` (e.g. you expected `float64` but got `object`) is a classic sign that a stray non-numeric value (a typo, a placeholder string like `"N/A"`) is silently corrupting the whole column's type — find and fix that value before modeling. `unexpected_column` isn't necessarily bad, just update your schema once you've confirmed it's intentional.
|
|
139
|
+
|
|
140
|
+
#### `range_validation(df, rules) -> DataFrame`
|
|
141
|
+
- **Checks:** numeric bounds and/or category whitelists, in one pass.
|
|
142
|
+
- **Expects:** `rules` — a dict of `{column_name: constraint_dict}`, where each constraint dict can have `"min"`, `"max"`, and/or `"allowed"` (a list of permitted values). Example: `{"age": {"min": 0, "max": 120}, "status": {"allowed": ["active", "closed", "pending"]}}`.
|
|
143
|
+
- **Returns:** one row per constraint you defined (a column with both `min` and `max` produces two rows), with `violation_count` and `violation_pct`.
|
|
144
|
+
- **Act on it:** any violation count above 0 needs a judgment call: is this bad data (a negative age → almost certainly an error, fix or drop it) or a rule that was wrong (maybe -1 is a legitimate "unknown" sentinel value in this dataset → adjust the rule, not the data). Don't silently clip or drop without checking which case you're in.
|
|
145
|
+
|
|
146
|
+
#### `cross_field_validation(df, rules) -> DataFrame`
|
|
147
|
+
- **Checks:** logical relationships *between* two columns in the same row.
|
|
148
|
+
- **Expects:** `rules` — a list of `(col_a, operator, col_b)` tuples. Supported operators: `>=`, `>`, `<=`, `<`, `==`, `!=`. Example: `[("end_date", ">=", "start_date"), ("net_income", "<=", "gross_income")]`. Rows where either column is missing are skipped, not flagged.
|
|
149
|
+
- **Returns:** one row per rule, with `violation_count` and `violation_pct`.
|
|
150
|
+
- **Act on it:** this is one of the highest-signal checks in the toolkit — a violation here usually means a genuine data entry or pipeline error (an end date before a start date isn't a modeling nuisance, it's a broken record). Pull the violating rows out and inspect them individually rather than aggregating past them.
|
|
151
|
+
|
|
152
|
+
### Master wrapper
|
|
153
|
+
|
|
154
|
+
#### `run_data_audit(df, target_col, id_col, expected_schema=None, range_rules=None, cross_field_rules=None) -> dict`
|
|
155
|
+
- **Checks:** everything above, in one call.
|
|
156
|
+
- **Expects:** `target_col` and `id_col` are required. `expected_schema`, `range_rules`, `cross_field_rules` are optional — pass them once you know what "valid" means for this dataset, and their reports get added to the output automatically.
|
|
157
|
+
- **Returns:** a dict keyed by report name (`"schema"`, `"target"`, `"id"`, `"duplicates"`, `"missingness"`, `"hidden_missingness"`, `"category_normalization"`, `"infinite_values"`, `"cardinality"`, `"constant_features"`, `"numeric_profile"`, plus the three validation reports if you supplied rules for them).
|
|
158
|
+
- **Act on it:** treat this as your standard "first thing I run on any new dataset" call. Loop through the dict and eyeball each report before writing a single line of feature engineering.
|
|
159
|
+
|
|
160
|
+
---
|
|
161
|
+
|
|
162
|
+
## `eda.py` reference
|
|
163
|
+
|
|
164
|
+
### Distributions
|
|
165
|
+
|
|
166
|
+
#### `target_distribution(df, target) -> DataFrame`
|
|
167
|
+
- **Checks:** the same thing as `target_report` in `audit.py` — count and percentage per class of the target.
|
|
168
|
+
- **Expects:** `target` is a column name.
|
|
169
|
+
- **Returns:** one row per class.
|
|
170
|
+
- **Act on it:** same as `target_report` — use it to catch class imbalance early. (These two functions currently duplicate each other; a future cleanup will likely have `eda` call `audit`'s version instead of recomputing it.)
|
|
171
|
+
|
|
172
|
+
#### `zero_percentage(df, columns) -> DataFrame`
|
|
173
|
+
- **Checks:** what fraction of each numeric column is exactly `0` (as opposed to missing).
|
|
174
|
+
- **Expects:** a list of numeric column names.
|
|
175
|
+
- **Returns:** `feature`, `zero_pct`, only for columns with at least one zero, sorted highest first.
|
|
176
|
+
- **Act on it:** a high zero percentage isn't automatically a problem — sometimes zero is a genuine value (zero purchases, zero late payments). But if it's unexpectedly high for a column like "income" or "age", it may actually be a placeholder for missing data that never got converted to `NaN`. Check the data dictionary or source system before treating zeros at face value.
|
|
177
|
+
|
|
178
|
+
### Outliers
|
|
179
|
+
|
|
180
|
+
#### `iqr_outlier_summary(df, columns) -> DataFrame`
|
|
181
|
+
- **Checks:** outliers per numeric column using the standard IQR rule (outside `[Q1 - 1.5×IQR, Q3 + 1.5×IQR]`).
|
|
182
|
+
- **Expects:** a list of numeric column names.
|
|
183
|
+
- **Returns:** `feature`, `q1`, `q3`, `iqr`, `lower_bound`, `upper_bound`, `outlier_count`, `outlier_pct`, sorted worst first.
|
|
184
|
+
- **Act on it:** a handful of outliers (under ~1-2%) is normal — look at a few individually to confirm they're real, not data entry errors, and decide whether to cap/winsorize, transform (log), or leave them (tree-based models are largely outlier-robust; linear/distance-based models are not). A very high outlier percentage usually means the IQR rule doesn't fit this column's distribution (e.g. it's naturally right-skewed, like income) rather than that 20% of your data is broken — pair this with `numeric_profile`'s skew column before reacting.
|
|
185
|
+
|
|
186
|
+
### Target relationships
|
|
187
|
+
|
|
188
|
+
#### `categorical_target_rate(df, feature, target) -> DataFrame`
|
|
189
|
+
- **Checks:** the target rate (e.g. default rate, churn rate) within each category of a feature.
|
|
190
|
+
- **Expects:** `target` should be numeric 0/1 (the function calls `.mean()` on it, which needs a numeric target).
|
|
191
|
+
- **Returns:** `feature` (category), `count`, `positives`, `target_rate`, `target_rate_pct`, sorted highest rate first.
|
|
192
|
+
- **Act on it:** big spread in target rate across categories → this feature likely has real predictive power, keep it. Also check `count` per category — a category with a dramatic rate but only 3 rows is noise, not signal; consider grouping small categories together.
|
|
193
|
+
|
|
194
|
+
#### `numeric_target_rate_by_quantile(df, feature, target, bins=10) -> DataFrame`
|
|
195
|
+
- **Checks:** the same idea as above, but for a continuous feature — splits it into quantile bins and shows the target rate per bin.
|
|
196
|
+
- **Expects:** `target` numeric 0/1. **Important:** the feature column must not contain `inf` values (see Known Quirks below) — run `infinite_value_report` first and clean those up.
|
|
197
|
+
- **Returns:** `bin` (the quantile range), `count`, `positives`, `target_rate`, `target_rate_pct`.
|
|
198
|
+
- **Act on it:** a target rate that rises or falls monotonically across bins → strong signal, this feature (or a transformed/binned version of it) will likely help a model. A non-monotonic, jagged pattern (e.g. high-low-high-low) → a linear model will miss this relationship even though it's real; a tree-based model or explicit binning will capture it better.
|
|
199
|
+
|
|
200
|
+
### Correlations
|
|
201
|
+
|
|
202
|
+
#### `correlation_matrix(df, columns, method="pearson") -> DataFrame`
|
|
203
|
+
- **Checks:** pairwise correlation between numeric columns.
|
|
204
|
+
- **Expects:** a list of numeric columns; `method` can be `"pearson"` (linear), `"spearman"` (monotonic/rank-based), or `"kendall"`.
|
|
205
|
+
- **Returns:** a square correlation matrix.
|
|
206
|
+
- **Act on it:** this is usually consumed by `high_correlation_pairs` and `plot_correlation_heatmap` rather than read raw — see below.
|
|
207
|
+
|
|
208
|
+
#### `high_correlation_pairs(corr_matrix, threshold=0.85) -> DataFrame`
|
|
209
|
+
- **Checks:** pulls out just the feature pairs from a correlation matrix that exceed a given absolute correlation.
|
|
210
|
+
- **Expects:** the output of `correlation_matrix(...)`.
|
|
211
|
+
- **Returns:** `feature_1`, `feature_2`, `correlation`, sorted by absolute strength. Empty DataFrame if nothing crosses the threshold.
|
|
212
|
+
- **Act on it:** pairs above ~0.85–0.9 are close to redundant. For linear/logistic regression, this kind of multicollinearity inflates coefficient variance and makes them unstable — drop one feature from each pair, or combine them (e.g. a ratio or PCA component). Tree-based models tolerate this better, but you're still paying for two features' worth of noise for one feature's worth of signal.
|
|
213
|
+
|
|
214
|
+
### Plots
|
|
215
|
+
|
|
216
|
+
All plotting functions call `plt.show()` and don't return a figure — they're meant for interactive/notebook use. Each one mirrors a data function above so you get the numbers and the picture together.
|
|
217
|
+
|
|
218
|
+
#### `plot_numeric_distribution(df, column)`
|
|
219
|
+
- **Shows:** a histogram (with KDE) and a boxplot side by side for one numeric column.
|
|
220
|
+
- **Use it after:** `numeric_profile` flags a column with high skew, or `iqr_outlier_summary` flags one with a high outlier percentage — this lets you see the shape, not just the number.
|
|
221
|
+
|
|
222
|
+
#### `plot_categorical_distribution(df, column, top_n=15)`
|
|
223
|
+
- **Shows:** a horizontal bar chart of category frequency for a categorical column (capped at the top N categories, default 15, so it stays readable on high-cardinality columns).
|
|
224
|
+
- **Use it after:** `cardinality_report` flags a column — visually confirm whether it's a few dominant categories with a long thin tail (a good candidate for "group rare into Other"), or genuinely spread out.
|
|
225
|
+
|
|
226
|
+
#### `plot_numeric_by_target(df, feature, target)`
|
|
227
|
+
- **Shows:** overlapping density curves of a numeric feature, split by target class.
|
|
228
|
+
- **Use it after:** you want to eyeball whether a feature separates the two classes at all, before formally checking with `numeric_target_rate_by_quantile`. Curves that barely overlap → strong feature. Curves that sit almost on top of each other → weak feature.
|
|
229
|
+
|
|
230
|
+
#### `plot_categorical_target_rate(df, feature, target)`
|
|
231
|
+
- **Shows:** the same information as `categorical_target_rate`, as a horizontal bar chart.
|
|
232
|
+
- **Use it after:** calling `categorical_target_rate` — bars make it much faster to spot which categories are driving risk/behavior than scanning a table.
|
|
233
|
+
|
|
234
|
+
#### `plot_numeric_target_rate(df, feature, target, bins=10)`
|
|
235
|
+
- **Shows:** the same information as `numeric_target_rate_by_quantile`, as a line plot across bins.
|
|
236
|
+
- **Use it after:** calling `numeric_target_rate_by_quantile` — the line makes monotonic vs. jagged relationships immediately visible (see the "Act on it" note above on why that distinction matters for model choice). Same `inf`-value caveat applies.
|
|
237
|
+
|
|
238
|
+
#### `plot_correlation_heatmap(corr_matrix)`
|
|
239
|
+
- **Shows:** a heatmap of a correlation matrix, annotated with the actual coefficient values.
|
|
240
|
+
- **Use it after:** calling `correlation_matrix` — faster than scanning `high_correlation_pairs` when you want the full picture of how everything relates to everything, not just the pairs above a threshold.
|
|
241
|
+
|
|
242
|
+
---
|
|
243
|
+
|
|
244
|
+
## Known quirks
|
|
245
|
+
|
|
246
|
+
- **`numeric_target_rate_by_quantile` (and its plot) will crash on `inf` values.** `pd.qcut` can't build bin edges around an infinite value. This is a real interaction between the two modules: `infinite_value_report` exists specifically to catch this kind of thing, so make it a habit to run it — and clean up whatever it finds — before reaching for any quantile-based EDA function.
|
|
247
|
+
- **`target_distribution` (eda.py) and `target_report` (audit.py) do the same calculation.** Harmless for now, but if you extend one, extend both, or better, consolidate them the next time you touch this code.
|
|
248
|
+
- **`schema_validation`'s dtype comparison is exact.** Depending on your pandas version, a plain string column may report as dtype `object` or as a newer string-specific dtype. If you get unexpected `dtype_mismatch` rows on columns that look fine, check what `df[col].dtype` actually prints in your environment and adjust `expected_schema` to match — this isn't a bug, it's the check doing its job of catching a dtype you didn't expect.
|
|
249
|
+
|
|
250
|
+
## Extending this toolkit
|
|
251
|
+
|
|
252
|
+
When adding a new function, keep the two existing conventions:
|
|
253
|
+
1. **Data functions return a DataFrame (or dict)**, never a plot. Keep plotting in separate `plot_*` functions that consume the data function's output where possible — that's what makes each check independently testable and reusable outside of a notebook.
|
|
254
|
+
2. **Document it here in the same format**: what it checks, what it expects, what it returns, and what decision to make from the result. A check nobody knows how to act on isn't pulling its weight in a toolkit like this one.
|
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
[build-system]
|
|
2
|
+
requires = ["setuptools>=61.0"]
|
|
3
|
+
build-backend = "setuptools.build_meta"
|
|
4
|
+
|
|
5
|
+
[project]
|
|
6
|
+
name = "ml-gearbox"
|
|
7
|
+
version = "0.1.0"
|
|
8
|
+
description = "Reusable Machine Learning Toolkit"
|
|
9
|
+
readme = "README.md"
|
|
10
|
+
requires-python = ">=3.10"
|
|
11
|
+
authors = [
|
|
12
|
+
{name = "Andrew Wamala Kyanjo"}
|
|
13
|
+
]
|
|
14
|
+
dependencies = [
|
|
15
|
+
"pandas",
|
|
16
|
+
"numpy",
|
|
17
|
+
"matplotlib",
|
|
18
|
+
"seaborn",
|
|
19
|
+
]
|
|
20
|
+
|
|
21
|
+
[project.optional-dependencies]
|
|
22
|
+
ml = ["scikit-learn"]
|
|
23
|
+
|
|
24
|
+
[project.urls]
|
|
25
|
+
Homepage = "https://github.com/AndrewKyanjo/ml-toolkit"
|
|
26
|
+
Issues = "https://github.com/AndrewKyanjo/ml-toolkit/issues"
|
|
27
|
+
|
|
28
|
+
[tool.setuptools.packages.find]
|
|
29
|
+
where = ["src"]
|
|
File without changes
|