goad-toolkit 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,231 @@
1
+ Metadata-Version: 2.4
2
+ Name: goad-toolkit
3
+ Version: 0.1.0
4
+ Summary: An extensible toolkit for Goal Oriented Analysis of Data
5
+ Author-email: raoul grouls <Raoul.Grouls@han.nl>
6
+ Requires-Python: >=3.12
7
+ Requires-Dist: loguru>=0.7.3
8
+ Requires-Dist: matplotlib>=3.10.1
9
+ Requires-Dist: numpy>=2.2.4
10
+ Requires-Dist: pandas>=2.2.3
11
+ Requires-Dist: pydantic>=2.10.6
12
+ Requires-Dist: requests>=2.32.3
13
+ Requires-Dist: scipy>=1.15.2
14
+ Requires-Dist: seaborn>=0.13.2
15
+ Requires-Dist: tqdm>=4.67.1
16
+ Description-Content-Type: text/markdown
17
+
18
+ # GOAD🐐 is the GOAT - Goal Oriented Analysis of Data
19
+ [![uv](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/uv/main/assets/badge/v0.json)](https://github.com/astral-sh/uv)
20
+
21
+ <p align="center">
22
+ <em>GOAD🐐 - When your data analysis is so fire🔥 it's got rizz✨</em>
23
+ </p>
24
+
25
+ GOAD🐐 is a flexible Python package for analyzing, transforming, and visualizing data with an emphasis on statistical distribution fitting and modular visualization components.
26
+
27
+ ## 📊 Features
28
+
29
+ - **Composable & extendable plotting system** - Build complex visualizations by combining simple components. You can extend the existing components with your own.
30
+ - **Statistical distribution fitting** - Automatically fit and compare distributions to your data. The distribution registry is extendable with additional distributions.
31
+ - **Extendable data transformation pipelines** - Chain and reuse data transformations into pipelines. Again, extendable with custom transformation components.
32
+
33
+ > Before GOAD🐐 : mid data
34
+ > After GOAD🐐 : data got infinity aura
35
+
36
+ ## 🚀 Quick Start
37
+
38
+ ### Installation
39
+ Using [uv](https://docs.astral.sh/uv/):
40
+ ```bash
41
+ uv install goad
42
+ ```
43
+
44
+ Or, if you prefer your dependencies to be installed 100x slower, with pip:
45
+ ```bash
46
+ pip install goad
47
+ ```
48
+
49
+ ## 📋 Demo: Linear Model Analysis
50
+
51
+ GOAD🐐 includes a comprehensive [demo](demo/linear.py) that shows how to use its components together.
52
+
53
+ ### Main capabilities
54
+ In the [demo/linear.py](demo/linear.py) file you can see a showcase of the main capabilities of GOAD🐐:
55
+ - create a data processing pipeline
56
+ - components are extendable, so you can easily add your own steps to a pipeline
57
+ - create visualisations by stacking components. `BasePlot` will handle boilerplate.
58
+ - the `DistributionFitter` will try to fit a few common distributions, and add statistical tests for you
59
+ - The results work together with the `visualizer.PlotFits` class to show the results
60
+
61
+ The main strenght of this module is not that these elements are there (even thought they are very useful). Its superpower is that everything is extendable: so you can use this as a start, and extend it with your own visualisations and analytics.
62
+
63
+ > POV: Your data just got GOADed🐐 and now it's giving main character energy
64
+
65
+ ## 📚 Core Components
66
+
67
+ #### 🔄 Extendable Data Transforms
68
+
69
+ GOAD🐐 provides a pipeline approach to transform your data:
70
+
71
+ ```python
72
+ from goad.datatransforms import Pipeline, ShiftValues, ZScaler
73
+
74
+ # Create a pipeline
75
+ pipeline = Pipeline()
76
+
77
+ # Add transformations
78
+ pipeline.add(ShiftValues, name="shift_deaths", column="deaths", period=-14)
79
+ pipeline.add(ZScaler, name="scale_tests", column="positivetests", rename=True)
80
+
81
+ # Apply all transformations
82
+ result = pipeline.apply(data)
83
+ ```
84
+
85
+ Available transforms include:
86
+ - `ShiftValues` - Shift values in a column by a specified period
87
+ - `DiffValues` - Calculate the difference between consecutive values
88
+ - `SelectDataRange` - Select rows within a specified date range
89
+ - `RollingAvg` - Calculate rolling average of a column
90
+ - `ZScaler` - Standardize values in a column
91
+
92
+ You can extend the pipeline with your own transformations by subclassing `BaseTransform`. The Zscaler is implemented as follows:
93
+
94
+ ```python
95
+ class ZScaler(TransformBase):
96
+ """Standardize the values in a column."""
97
+ def transform(
98
+ self, data: pd.DataFrame, column: str, rename: bool = False
99
+ ) -> pd.DataFrame:
100
+ """Standardize the values in a column."""
101
+ if rename:
102
+ colname = f"{column}_zscore"
103
+ else:
104
+ colname = column
105
+ data[colname] = (data[column] - data[column].mean()) / data[column].std()
106
+ return data
107
+ ```
108
+
109
+ ### 📊 Visualization System
110
+
111
+ GOAD🐐 visualization system is built on a composable architecture that allows you to build complex plots by combining simpler components:
112
+
113
+ ```python
114
+ from goad.visualizer import PlotSettings, ResidualPlot
115
+
116
+ # Create plot settings
117
+ plotsettings = PlotSettings(
118
+ xlabel="date",
119
+ ylabel="normalized values",
120
+ title="Z-Scores of Deaths and Positive Tests",
121
+ )
122
+
123
+ class LinePlot(BasePlot):
124
+ """Plot a line plot using seaborn."""
125
+ def build(self, data: pd.DataFrame, **kwargs):
126
+ sns.lineplot(data=data, ax=self.ax, **kwargs)
127
+ return self.fig, self.ax
128
+
129
+
130
+ class ComparePlot(BasePlot):
131
+ def build(self, data: pd.DataFrame, x: str, y1: str, y2: str, **kwargs):
132
+ compare = LinePlot(self.settings)
133
+ self.plot_on(compare, data=data, x=x, y=y1, label=y1, **kwargs)
134
+ self.plot_on(compare, data=data, x=x, y=y2, label=y2, **kwargs)
135
+ plt.xticks(rotation=45)
136
+
137
+ return self.fig, self.ax
138
+
139
+ compareplot = ComparePlot(plotsettings)
140
+ compareplot.plot(
141
+ data=data, x="date", y1="deaths_shifted_zscore", y2="positivetests_zscore"
142
+ )
143
+ ```
144
+ ![zscore](img/zscores.png)
145
+ This extendable strategy lets BasePlot handle the boilerplate, while you can focus on creating the visualizations you need.
146
+ It is also easier to reuse components in different contexts.
147
+ ### 📈 Distribution Fitting
148
+
149
+ GOAD🐐 includes tools for fitting statistical distributions to your data:
150
+
151
+ ```python
152
+ from goad.analytics import DistributionFitter
153
+ from goad.visualizer import PlotSettings, FitPlotSettings, PlotFits
154
+
155
+ fitter = DistributionFitter()
156
+ fits = fitter.fit(data["residual"], discrete=False) # we have to decide if the data is discrete or not
157
+ best = fitter.best(fits)
158
+ settings = PlotSettings(
159
+ figsize=(12, 6), title="Residuals", xlabel="error", ylabel="probability"
160
+ )
161
+ fitplotsettings = FitPlotSettings(bins=30, max_fits=3)
162
+ fitplotter = PlotFits(settings)
163
+ fig = fitplotter.plot(
164
+ data=data["residual"], fit_results=fits, fitplotsettings=fitplotsettings
165
+ )
166
+ ```
167
+ For the [kstest](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.kstest.html), the null hypothesis is that the two distributions are identical. In this example, the p-values are below 0.05, so we can reject the null hypothesis and conclude that the data does not follow any of these.
168
+
169
+ The plots are sorted by log-likelihood, which means there is no good fit with a distribution in this case.
170
+ ![residuals](img/distribution_fit.png)
171
+
172
+ ### 🧩 Extending with Custom Distributions
173
+
174
+ You can easily register new distributions:
175
+
176
+ ```python
177
+ from goad.distributions import DistributionRegistry
178
+ from scipy import stats
179
+
180
+ # Create registry
181
+ registry = DistributionRegistry()
182
+
183
+ # Register a new distribution
184
+ registry.register_distribution(
185
+ name="negative_binomial",
186
+ dist=stats.nbinom,
187
+ is_discrete=True,
188
+ num_params=2
189
+ )
190
+
191
+ # Now it will be used automatically in the DistributionFitter for discrete fits
192
+ from goad.analytics import DistributionFitter
193
+ fitter = DistributionFitter()
194
+ print(fitter.registry) # shows all registered distributions
195
+ ```
196
+
197
+
198
+
199
+ ## 🔧 Advanced Usage: Composing Plots
200
+
201
+ GOAD🐐 has a powerful plotting system that allows you to combine plot elements:
202
+
203
+ ```python
204
+ from goad.visualizer import BasePlot, LinePlot, BarWithDates, VerticalDate
205
+
206
+ # Use a base plot to create a composite
207
+ class MyCompositePlot(BasePlot):
208
+ def build(self, data: pd.DataFrame, x: str, y1: str, y2: str, special_date: str):
209
+ # Plot the first component - a line plot
210
+ line_plot = LinePlot(self.settings)
211
+ self.plot_on(line_plot, data=data, x=x, y=y1, label=y1)
212
+
213
+ # Plot the second component - a bar chart
214
+ bar_plot = BarWithDates(self.settings)
215
+ self.plot_on(bar_plot, data=data, x=x, y=y2)
216
+
217
+ # Add a vertical line
218
+ vline = VerticalDate(self.settings)
219
+ self.plot_on(vline, date=special_date, label="Important Event")
220
+ return self.fig, self.ax
221
+ ```
222
+
223
+ ## 🤝 Contributing
224
+
225
+ Contributions are welcome! Please feel free to submit a Pull Request.
226
+
227
+ ---
228
+
229
+ <p align="center">
230
+ <em>GOAD🐐 - When your data analysis is so fire🔥 it's got rizz✨</em>
231
+ </p>
@@ -0,0 +1,14 @@
1
+ goad_toolkit/__init__.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
2
+ goad_toolkit/analytics.py,sha256=X-HJkd2nKwxyZ1xHZ_LS5qqECp6vO3KiQZVDR54mFRE,11785
3
+ goad_toolkit/config.py,sha256=pILb5LyysMBMhMcuJSEgmSWyqevb3UcSNmTMgA7kR1o,443
4
+ goad_toolkit/dataprocessor.py,sha256=jheEOTtqkEwMQ0zN3fryqe7mMUyUCzIcSQLodNeDBws,1660
5
+ goad_toolkit/datatransforms.py,sha256=KDfI3AbmfAVl51krIymfqH_StHkWf-Buja3RPOS-z38,6222
6
+ goad_toolkit/distributions.py,sha256=hZfKHVmGe7W1XiCTp7X_Zm2gLSLMQbecPxQb4Ch4j2c,2916
7
+ goad_toolkit/filehandler.py,sha256=Xz-9-Cqs3oPwIsWDNNosM1skTn6njvLGIlF0osPPC9A,2125
8
+ goad_toolkit/main.py,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
9
+ goad_toolkit/models.py,sha256=jLsFOFarJsizXST8X9LJytehwxq-KTU4WkeEM-MOGYM,1015
10
+ goad_toolkit/visualizer.py,sha256=-uhdM08CHqfuK0GfkAvThfsG1_IcfiiBMZpf8DaE4-Q,16389
11
+ goad_toolkit-0.1.0.dist-info/METADATA,sha256=TJ8xC_BXR70mCMN00-0X0qwYiVnDezaT8UDJvvRWkHQ,8332
12
+ goad_toolkit-0.1.0.dist-info/WHEEL,sha256=qtCwoSJWgHk21S1Kb4ihdzI2rlJ1ZKaIurTj_ngOhyQ,87
13
+ goad_toolkit-0.1.0.dist-info/entry_points.txt,sha256=m9nMY1WMdFknEkSfuq_jh64X1AL9M_U5Fy0UrqOnm10,35
14
+ goad_toolkit-0.1.0.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.27.0
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ goad = goad:main