mlarena 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
mlarena-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2024 Your Name
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
mlarena-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,136 @@
1
+ Metadata-Version: 2.3
2
+ Name: mlarena
3
+ Version: 0.1.0
4
+ Summary: An algorithm-agnostic machine learning toolkit for model training, diagnostics and optimization
5
+ License: MIT
6
+ Keywords: machine-learning,data-science,preprocessing,pipeline
7
+ Author: Mena Wang
8
+ Author-email: ningwang25@gmail.com
9
+ Requires-Python: >=3.10
10
+ Classifier: Development Status :: 3 - Alpha
11
+ Classifier: Intended Audience :: Science/Research
12
+ Classifier: License :: OSI Approved :: MIT License
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
19
+ Provides-Extra: demo-dependencies
20
+ Requires-Dist: hyperopt (>=0.2.7)
21
+ Requires-Dist: matplotlib (>=3.4.0)
22
+ Requires-Dist: mlflow (>=2.0.0)
23
+ Requires-Dist: numpy (>=1.21.0)
24
+ Requires-Dist: pandas (>=1.3.0)
25
+ Requires-Dist: scikit-learn (>=0.24.0)
26
+ Requires-Dist: shap (>=0.41.0)
27
+ Project-URL: Homepage, https://github.com/MenaWANG/mlarena
28
+ Project-URL: Repository, https://github.com/MenaWANG/mlarena
29
+ Description-Content-Type: text/markdown
30
+
31
+ # MLArena
32
+
33
+ An algorithm-agnostic machine learning toolkit for model training, diagnostics and optimization.
34
+
35
+ ## Features
36
+
37
+ - **Comprehensive ML Pipeline**:
38
+ - End-to-end workflow from preprocessing to deployment
39
+ - Model-agnostic design (works with any scikit-learn compatible model)
40
+ - Support for both classification and regression tasks
41
+ - Early stopping and validation set support
42
+ - MLflow integration for experiment tracking and deployment
43
+
44
+ - **Intelligent Preprocessing**:
45
+ - Automated feature type detection and handling
46
+ - Smart encoding recommendations based on feature cardinality and rare category
47
+ - Target encoding with visualization to support smoothing parameter selection
48
+ - Missing value handling with configurable strategies
49
+ - Feature selection recommendations with mutual information analysis
50
+
51
+ - **Advanced Model Evaluation**:
52
+ - Comprehensive metrics for both classification and regression
53
+ - Diagnostic visualization of model performance
54
+ - Threshold analysis for classification tasks
55
+ - SHAP-based model explanations (global and local)
56
+ - Cross-validation with variance penalty
57
+
58
+ - **Hyperparameter Optimization**:
59
+ - Bayesian optimization with Hyperopt
60
+ - Cross-validation based tuning
61
+ - Parallel coordinates visualization for search space analysis
62
+ - Early stopping to prevent overfitting
63
+ - Variance penalty to ensure stable solutions
64
+
65
+
66
+ ## Installation
67
+
68
+ ```bash
69
+ pip install mlarena
70
+ ```
71
+
72
+ ## Quick Start
73
+
74
+ ```python
75
+ from mlarena import PreProcessor, ML_PIPELINE
76
+ from sklearn.ensemble import RandomForestClassifier
77
+
78
+ # Initialize the preprocessor
79
+ preprocessor = PreProcessor(
80
+ num_impute_strategy='median',
81
+ cat_impute_strategy='most_frequent'
82
+ )
83
+
84
+ # Initialize the pipeline
85
+ ml_pipeline = ML_PIPELINE(
86
+ model = RandomForestClassifier(),
87
+ preprocessor = preprocessor
88
+ )
89
+
90
+ # Train the model
91
+ ml_pipeline.fit(X_train, y_train)
92
+
93
+ # Make predictions
94
+ y_pred = ml_pipeline.predict(X_test)
95
+
96
+ # Comprehensive Evaluation Report and Visuals
97
+ results = ml_pipeline.evaluate(X_test, y_test)
98
+
99
+ # Explain the model
100
+ ml_pipeline.explain_model(X_test)
101
+
102
+ ```
103
+
104
+ ## Documentation
105
+
106
+ ### PreProcessor
107
+
108
+ The `PreProcessor` class handles all data preprocessing tasks:
109
+
110
+ - Filter Feature Selection
111
+ - Categorical encoding (OneHot, Target)
112
+ - Recommendation of encoding strategy
113
+ - Plot to compare target encoding smoothing parameters
114
+ - Numeric scaling
115
+ - Missing value imputation
116
+
117
+ ### ML_PIPELINE
118
+
119
+ The `ML_PIPELINE` class provides a complete machine learning workflow:
120
+
121
+ - Algorithm agnostic model wrapper
122
+ - Support both classification (binary) and regression algorithms
123
+ - Model training and scoring
124
+ - Model global and local explanation
125
+ - Model evaluation with comprehensive reporting and plots
126
+ - Iterative hyperparameter tuning with diagnostic plot
127
+ - Threshold analysis and optimization for classification models
128
+
129
+
130
+ ## Contributing
131
+
132
+ Contributions are welcome! Please feel free to submit a Pull Request.
133
+
134
+ ## License
135
+
136
+ This project is licensed under the MIT License - see the LICENSE file for details.
@@ -0,0 +1,106 @@
1
+ # MLArena
2
+
3
+ An algorithm-agnostic machine learning toolkit for model training, diagnostics and optimization.
4
+
5
+ ## Features
6
+
7
+ - **Comprehensive ML Pipeline**:
8
+ - End-to-end workflow from preprocessing to deployment
9
+ - Model-agnostic design (works with any scikit-learn compatible model)
10
+ - Support for both classification and regression tasks
11
+ - Early stopping and validation set support
12
+ - MLflow integration for experiment tracking and deployment
13
+
14
+ - **Intelligent Preprocessing**:
15
+ - Automated feature type detection and handling
16
+ - Smart encoding recommendations based on feature cardinality and rare category
17
+ - Target encoding with visualization to support smoothing parameter selection
18
+ - Missing value handling with configurable strategies
19
+ - Feature selection recommendations with mutual information analysis
20
+
21
+ - **Advanced Model Evaluation**:
22
+ - Comprehensive metrics for both classification and regression
23
+ - Diagnostic visualization of model performance
24
+ - Threshold analysis for classification tasks
25
+ - SHAP-based model explanations (global and local)
26
+ - Cross-validation with variance penalty
27
+
28
+ - **Hyperparameter Optimization**:
29
+ - Bayesian optimization with Hyperopt
30
+ - Cross-validation based tuning
31
+ - Parallel coordinates visualization for search space analysis
32
+ - Early stopping to prevent overfitting
33
+ - Variance penalty to ensure stable solutions
34
+
35
+
36
+ ## Installation
37
+
38
+ ```bash
39
+ pip install mlarena
40
+ ```
41
+
42
+ ## Quick Start
43
+
44
+ ```python
45
+ from mlarena import PreProcessor, ML_PIPELINE
46
+ from sklearn.ensemble import RandomForestClassifier
47
+
48
+ # Initialize the preprocessor
49
+ preprocessor = PreProcessor(
50
+ num_impute_strategy='median',
51
+ cat_impute_strategy='most_frequent'
52
+ )
53
+
54
+ # Initialize the pipeline
55
+ ml_pipeline = ML_PIPELINE(
56
+ model = RandomForestClassifier(),
57
+ preprocessor = preprocessor
58
+ )
59
+
60
+ # Train the model
61
+ ml_pipeline.fit(X_train, y_train)
62
+
63
+ # Make predictions
64
+ y_pred = ml_pipeline.predict(X_test)
65
+
66
+ # Comprehensive Evaluation Report and Visuals
67
+ results = ml_pipeline.evaluate(X_test, y_test)
68
+
69
+ # Explain the model
70
+ ml_pipeline.explain_model(X_test)
71
+
72
+ ```
73
+
74
+ ## Documentation
75
+
76
+ ### PreProcessor
77
+
78
+ The `PreProcessor` class handles all data preprocessing tasks:
79
+
80
+ - Filter Feature Selection
81
+ - Categorical encoding (OneHot, Target)
82
+ - Recommendation of encoding strategy
83
+ - Plot to compare target encoding smoothing parameters
84
+ - Numeric scaling
85
+ - Missing value imputation
86
+
87
+ ### ML_PIPELINE
88
+
89
+ The `ML_PIPELINE` class provides a complete machine learning workflow:
90
+
91
+ - Algorithm agnostic model wrapper
92
+ - Support both classification (binary) and regression algorithms
93
+ - Model training and scoring
94
+ - Model global and local explanation
95
+ - Model evaluation with comprehensive reporting and plots
96
+ - Iterative hyperparameter tuning with diagnostic plot
97
+ - Threshold analysis and optimization for classification models
98
+
99
+
100
+ ## Contributing
101
+
102
+ Contributions are welcome! Please feel free to submit a Pull Request.
103
+
104
+ ## License
105
+
106
+ This project is licensed under the MIT License - see the LICENSE file for details.
@@ -0,0 +1,13 @@
1
+ """
2
+ MLArena - A comprehensive ML pipeline wrapper for scikit-learn compatible models.
3
+
4
+ This package provides:
5
+ - PreProcessor: Advanced data preprocessing with feature analysis and smart encoding
6
+ - ML_PIPELINE: End-to-end ML pipeline with model training, evaluation, and deployment
7
+ """
8
+
9
+ from .preprocessor import PreProcessor
10
+ from .pipeline import ML_PIPELINE
11
+
12
+ __version__ = "0.1.0"
13
+ __all__ = ["PreProcessor", "ML_PIPELINE"]