goad-toolkit 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- goad_toolkit/__init__.py +0 -0
- goad_toolkit/analytics.py +325 -0
- goad_toolkit/config.py +18 -0
- goad_toolkit/dataprocessor.py +51 -0
- goad_toolkit/datatransforms.py +184 -0
- goad_toolkit/distributions.py +80 -0
- goad_toolkit/filehandler.py +65 -0
- goad_toolkit/main.py +0 -0
- goad_toolkit/models.py +44 -0
- goad_toolkit/visualizer.py +505 -0
- goad_toolkit-0.1.0.dist-info/METADATA +231 -0
- goad_toolkit-0.1.0.dist-info/RECORD +14 -0
- goad_toolkit-0.1.0.dist-info/WHEEL +4 -0
- goad_toolkit-0.1.0.dist-info/entry_points.txt +2 -0
|
@@ -0,0 +1,231 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: goad-toolkit
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: An extensible toolkit for Goal Oriented Analysis of Data
|
|
5
|
+
Author-email: raoul grouls <Raoul.Grouls@han.nl>
|
|
6
|
+
Requires-Python: >=3.12
|
|
7
|
+
Requires-Dist: loguru>=0.7.3
|
|
8
|
+
Requires-Dist: matplotlib>=3.10.1
|
|
9
|
+
Requires-Dist: numpy>=2.2.4
|
|
10
|
+
Requires-Dist: pandas>=2.2.3
|
|
11
|
+
Requires-Dist: pydantic>=2.10.6
|
|
12
|
+
Requires-Dist: requests>=2.32.3
|
|
13
|
+
Requires-Dist: scipy>=1.15.2
|
|
14
|
+
Requires-Dist: seaborn>=0.13.2
|
|
15
|
+
Requires-Dist: tqdm>=4.67.1
|
|
16
|
+
Description-Content-Type: text/markdown
|
|
17
|
+
|
|
18
|
+
# GOAD🐐 is the GOAT - Goal Oriented Analysis of Data
|
|
19
|
+
[](https://github.com/astral-sh/uv)
|
|
20
|
+
|
|
21
|
+
<p align="center">
|
|
22
|
+
<em>GOAD🐐 - When your data analysis is so fire🔥 it's got rizz✨</em>
|
|
23
|
+
</p>
|
|
24
|
+
|
|
25
|
+
GOAD🐐 is a flexible Python package for analyzing, transforming, and visualizing data with an emphasis on statistical distribution fitting and modular visualization components.
|
|
26
|
+
|
|
27
|
+
## 📊 Features
|
|
28
|
+
|
|
29
|
+
- **Composable & extendable plotting system** - Build complex visualizations by combining simple components. You can extend the existing components with your own.
|
|
30
|
+
- **Statistical distribution fitting** - Automatically fit and compare distributions to your data. The distribution registry is extendable with additional distributions.
|
|
31
|
+
- **Extendable data transformation pipelines** - Chain and reuse data transformations into pipelines. Again, extendable with custom transformation components.
|
|
32
|
+
|
|
33
|
+
> Before GOAD🐐 : mid data
|
|
34
|
+
> After GOAD🐐 : data got infinity aura
|
|
35
|
+
|
|
36
|
+
## 🚀 Quick Start
|
|
37
|
+
|
|
38
|
+
### Installation
|
|
39
|
+
Using [uv](https://docs.astral.sh/uv/):
|
|
40
|
+
```bash
|
|
41
|
+
uv install goad
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Or, if you prefer your dependencies to be installed 100x slower, with pip:
|
|
45
|
+
```bash
|
|
46
|
+
pip install goad
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
## 📋 Demo: Linear Model Analysis
|
|
50
|
+
|
|
51
|
+
GOAD🐐 includes a comprehensive [demo](demo/linear.py) that shows how to use its components together.
|
|
52
|
+
|
|
53
|
+
### Main capabilities
|
|
54
|
+
In the [demo/linear.py](demo/linear.py) file you can see a showcase of the main capabilities of GOAD🐐:
|
|
55
|
+
- create a data processing pipeline
|
|
56
|
+
- components are extendable, so you can easily add your own steps to a pipeline
|
|
57
|
+
- create visualisations by stacking components. `BasePlot` will handle boilerplate.
|
|
58
|
+
- the `DistributionFitter` will try to fit a few common distributions, and add statistical tests for you
|
|
59
|
+
- The results work together with the `visualizer.PlotFits` class to show the results
|
|
60
|
+
|
|
61
|
+
The main strenght of this module is not that these elements are there (even thought they are very useful). Its superpower is that everything is extendable: so you can use this as a start, and extend it with your own visualisations and analytics.
|
|
62
|
+
|
|
63
|
+
> POV: Your data just got GOADed🐐 and now it's giving main character energy
|
|
64
|
+
|
|
65
|
+
## 📚 Core Components
|
|
66
|
+
|
|
67
|
+
#### 🔄 Extendable Data Transforms
|
|
68
|
+
|
|
69
|
+
GOAD🐐 provides a pipeline approach to transform your data:
|
|
70
|
+
|
|
71
|
+
```python
|
|
72
|
+
from goad.datatransforms import Pipeline, ShiftValues, ZScaler
|
|
73
|
+
|
|
74
|
+
# Create a pipeline
|
|
75
|
+
pipeline = Pipeline()
|
|
76
|
+
|
|
77
|
+
# Add transformations
|
|
78
|
+
pipeline.add(ShiftValues, name="shift_deaths", column="deaths", period=-14)
|
|
79
|
+
pipeline.add(ZScaler, name="scale_tests", column="positivetests", rename=True)
|
|
80
|
+
|
|
81
|
+
# Apply all transformations
|
|
82
|
+
result = pipeline.apply(data)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Available transforms include:
|
|
86
|
+
- `ShiftValues` - Shift values in a column by a specified period
|
|
87
|
+
- `DiffValues` - Calculate the difference between consecutive values
|
|
88
|
+
- `SelectDataRange` - Select rows within a specified date range
|
|
89
|
+
- `RollingAvg` - Calculate rolling average of a column
|
|
90
|
+
- `ZScaler` - Standardize values in a column
|
|
91
|
+
|
|
92
|
+
You can extend the pipeline with your own transformations by subclassing `BaseTransform`. The Zscaler is implemented as follows:
|
|
93
|
+
|
|
94
|
+
```python
|
|
95
|
+
class ZScaler(TransformBase):
|
|
96
|
+
"""Standardize the values in a column."""
|
|
97
|
+
def transform(
|
|
98
|
+
self, data: pd.DataFrame, column: str, rename: bool = False
|
|
99
|
+
) -> pd.DataFrame:
|
|
100
|
+
"""Standardize the values in a column."""
|
|
101
|
+
if rename:
|
|
102
|
+
colname = f"{column}_zscore"
|
|
103
|
+
else:
|
|
104
|
+
colname = column
|
|
105
|
+
data[colname] = (data[column] - data[column].mean()) / data[column].std()
|
|
106
|
+
return data
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
### 📊 Visualization System
|
|
110
|
+
|
|
111
|
+
GOAD🐐 visualization system is built on a composable architecture that allows you to build complex plots by combining simpler components:
|
|
112
|
+
|
|
113
|
+
```python
|
|
114
|
+
from goad.visualizer import PlotSettings, ResidualPlot
|
|
115
|
+
|
|
116
|
+
# Create plot settings
|
|
117
|
+
plotsettings = PlotSettings(
|
|
118
|
+
xlabel="date",
|
|
119
|
+
ylabel="normalized values",
|
|
120
|
+
title="Z-Scores of Deaths and Positive Tests",
|
|
121
|
+
)
|
|
122
|
+
|
|
123
|
+
class LinePlot(BasePlot):
|
|
124
|
+
"""Plot a line plot using seaborn."""
|
|
125
|
+
def build(self, data: pd.DataFrame, **kwargs):
|
|
126
|
+
sns.lineplot(data=data, ax=self.ax, **kwargs)
|
|
127
|
+
return self.fig, self.ax
|
|
128
|
+
|
|
129
|
+
|
|
130
|
+
class ComparePlot(BasePlot):
|
|
131
|
+
def build(self, data: pd.DataFrame, x: str, y1: str, y2: str, **kwargs):
|
|
132
|
+
compare = LinePlot(self.settings)
|
|
133
|
+
self.plot_on(compare, data=data, x=x, y=y1, label=y1, **kwargs)
|
|
134
|
+
self.plot_on(compare, data=data, x=x, y=y2, label=y2, **kwargs)
|
|
135
|
+
plt.xticks(rotation=45)
|
|
136
|
+
|
|
137
|
+
return self.fig, self.ax
|
|
138
|
+
|
|
139
|
+
compareplot = ComparePlot(plotsettings)
|
|
140
|
+
compareplot.plot(
|
|
141
|
+
data=data, x="date", y1="deaths_shifted_zscore", y2="positivetests_zscore"
|
|
142
|
+
)
|
|
143
|
+
```
|
|
144
|
+

|
|
145
|
+
This extendable strategy lets BasePlot handle the boilerplate, while you can focus on creating the visualizations you need.
|
|
146
|
+
It is also easier to reuse components in different contexts.
|
|
147
|
+
### 📈 Distribution Fitting
|
|
148
|
+
|
|
149
|
+
GOAD🐐 includes tools for fitting statistical distributions to your data:
|
|
150
|
+
|
|
151
|
+
```python
|
|
152
|
+
from goad.analytics import DistributionFitter
|
|
153
|
+
from goad.visualizer import PlotSettings, FitPlotSettings, PlotFits
|
|
154
|
+
|
|
155
|
+
fitter = DistributionFitter()
|
|
156
|
+
fits = fitter.fit(data["residual"], discrete=False) # we have to decide if the data is discrete or not
|
|
157
|
+
best = fitter.best(fits)
|
|
158
|
+
settings = PlotSettings(
|
|
159
|
+
figsize=(12, 6), title="Residuals", xlabel="error", ylabel="probability"
|
|
160
|
+
)
|
|
161
|
+
fitplotsettings = FitPlotSettings(bins=30, max_fits=3)
|
|
162
|
+
fitplotter = PlotFits(settings)
|
|
163
|
+
fig = fitplotter.plot(
|
|
164
|
+
data=data["residual"], fit_results=fits, fitplotsettings=fitplotsettings
|
|
165
|
+
)
|
|
166
|
+
```
|
|
167
|
+
For the [kstest](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.kstest.html), the null hypothesis is that the two distributions are identical. In this example, the p-values are below 0.05, so we can reject the null hypothesis and conclude that the data does not follow any of these.
|
|
168
|
+
|
|
169
|
+
The plots are sorted by log-likelihood, which means there is no good fit with a distribution in this case.
|
|
170
|
+

|
|
171
|
+
|
|
172
|
+
### 🧩 Extending with Custom Distributions
|
|
173
|
+
|
|
174
|
+
You can easily register new distributions:
|
|
175
|
+
|
|
176
|
+
```python
|
|
177
|
+
from goad.distributions import DistributionRegistry
|
|
178
|
+
from scipy import stats
|
|
179
|
+
|
|
180
|
+
# Create registry
|
|
181
|
+
registry = DistributionRegistry()
|
|
182
|
+
|
|
183
|
+
# Register a new distribution
|
|
184
|
+
registry.register_distribution(
|
|
185
|
+
name="negative_binomial",
|
|
186
|
+
dist=stats.nbinom,
|
|
187
|
+
is_discrete=True,
|
|
188
|
+
num_params=2
|
|
189
|
+
)
|
|
190
|
+
|
|
191
|
+
# Now it will be used automatically in the DistributionFitter for discrete fits
|
|
192
|
+
from goad.analytics import DistributionFitter
|
|
193
|
+
fitter = DistributionFitter()
|
|
194
|
+
print(fitter.registry) # shows all registered distributions
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
|
|
198
|
+
|
|
199
|
+
## 🔧 Advanced Usage: Composing Plots
|
|
200
|
+
|
|
201
|
+
GOAD🐐 has a powerful plotting system that allows you to combine plot elements:
|
|
202
|
+
|
|
203
|
+
```python
|
|
204
|
+
from goad.visualizer import BasePlot, LinePlot, BarWithDates, VerticalDate
|
|
205
|
+
|
|
206
|
+
# Use a base plot to create a composite
|
|
207
|
+
class MyCompositePlot(BasePlot):
|
|
208
|
+
def build(self, data: pd.DataFrame, x: str, y1: str, y2: str, special_date: str):
|
|
209
|
+
# Plot the first component - a line plot
|
|
210
|
+
line_plot = LinePlot(self.settings)
|
|
211
|
+
self.plot_on(line_plot, data=data, x=x, y=y1, label=y1)
|
|
212
|
+
|
|
213
|
+
# Plot the second component - a bar chart
|
|
214
|
+
bar_plot = BarWithDates(self.settings)
|
|
215
|
+
self.plot_on(bar_plot, data=data, x=x, y=y2)
|
|
216
|
+
|
|
217
|
+
# Add a vertical line
|
|
218
|
+
vline = VerticalDate(self.settings)
|
|
219
|
+
self.plot_on(vline, date=special_date, label="Important Event")
|
|
220
|
+
return self.fig, self.ax
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
## 🤝 Contributing
|
|
224
|
+
|
|
225
|
+
Contributions are welcome! Please feel free to submit a Pull Request.
|
|
226
|
+
|
|
227
|
+
---
|
|
228
|
+
|
|
229
|
+
<p align="center">
|
|
230
|
+
<em>GOAD🐐 - When your data analysis is so fire🔥 it's got rizz✨</em>
|
|
231
|
+
</p>
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
goad_toolkit/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
2
|
+
goad_toolkit/analytics.py,sha256=X-HJkd2nKwxyZ1xHZ_LS5qqECp6vO3KiQZVDR54mFRE,11785
|
|
3
|
+
goad_toolkit/config.py,sha256=pILb5LyysMBMhMcuJSEgmSWyqevb3UcSNmTMgA7kR1o,443
|
|
4
|
+
goad_toolkit/dataprocessor.py,sha256=jheEOTtqkEwMQ0zN3fryqe7mMUyUCzIcSQLodNeDBws,1660
|
|
5
|
+
goad_toolkit/datatransforms.py,sha256=KDfI3AbmfAVl51krIymfqH_StHkWf-Buja3RPOS-z38,6222
|
|
6
|
+
goad_toolkit/distributions.py,sha256=hZfKHVmGe7W1XiCTp7X_Zm2gLSLMQbecPxQb4Ch4j2c,2916
|
|
7
|
+
goad_toolkit/filehandler.py,sha256=Xz-9-Cqs3oPwIsWDNNosM1skTn6njvLGIlF0osPPC9A,2125
|
|
8
|
+
goad_toolkit/main.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
9
|
+
goad_toolkit/models.py,sha256=jLsFOFarJsizXST8X9LJytehwxq-KTU4WkeEM-MOGYM,1015
|
|
10
|
+
goad_toolkit/visualizer.py,sha256=-uhdM08CHqfuK0GfkAvThfsG1_IcfiiBMZpf8DaE4-Q,16389
|
|
11
|
+
goad_toolkit-0.1.0.dist-info/METADATA,sha256=TJ8xC_BXR70mCMN00-0X0qwYiVnDezaT8UDJvvRWkHQ,8332
|
|
12
|
+
goad_toolkit-0.1.0.dist-info/WHEEL,sha256=qtCwoSJWgHk21S1Kb4ihdzI2rlJ1ZKaIurTj_ngOhyQ,87
|
|
13
|
+
goad_toolkit-0.1.0.dist-info/entry_points.txt,sha256=m9nMY1WMdFknEkSfuq_jh64X1AL9M_U5Fy0UrqOnm10,35
|
|
14
|
+
goad_toolkit-0.1.0.dist-info/RECORD,,
|