binboost 0.2.1__tar.gz → 0.2.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- binboost-0.2.3/PKG-INFO +129 -0
- binboost-0.2.3/README.md +108 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost/__init__.py +2 -1
- {binboost-0.2.1 → binboost-0.2.3}/binboost/binarizer.py +1 -1
- {binboost-0.2.1 → binboost-0.2.3}/binboost/binboost.py +42 -20
- {binboost-0.2.1 → binboost-0.2.3}/binboost/loss.py +1 -1
- binboost-0.2.3/binboost/multiclass.py +78 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost/rule.py +78 -76
- binboost-0.2.3/binboost.egg-info/PKG-INFO +129 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost.egg-info/SOURCES.txt +1 -0
- {binboost-0.2.1 → binboost-0.2.3}/pyproject.toml +1 -1
- {binboost-0.2.1 → binboost-0.2.3}/setup.py +1 -1
- binboost-0.2.1/PKG-INFO +0 -101
- binboost-0.2.1/README.md +0 -80
- binboost-0.2.1/binboost.egg-info/PKG-INFO +0 -101
- {binboost-0.2.1 → binboost-0.2.3}/LICENSE +0 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost.egg-info/dependency_links.txt +0 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost.egg-info/requires.txt +0 -0
- {binboost-0.2.1 → binboost-0.2.3}/binboost.egg-info/top_level.txt +0 -0
- {binboost-0.2.1 → binboost-0.2.3}/setup.cfg +0 -0
binboost-0.2.3/PKG-INFO
ADDED
|
@@ -0,0 +1,129 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: binboost
|
|
3
|
+
Version: 0.2.3
|
|
4
|
+
Summary: Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner
|
|
5
|
+
Author: Rangga Wahyu Pratama
|
|
6
|
+
License: MIT
|
|
7
|
+
Keywords: gradient boosting,logical rules,binary features,interpretable machine learning
|
|
8
|
+
Classifier: Programming Language :: Python :: 3
|
|
9
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
+
Classifier: Operating System :: OS Independent
|
|
11
|
+
Classifier: Intended Audience :: Science/Research
|
|
12
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
13
|
+
Requires-Python: >=3.8
|
|
14
|
+
Description-Content-Type: text/markdown
|
|
15
|
+
License-File: LICENSE
|
|
16
|
+
Requires-Dist: numpy>=1.21.0
|
|
17
|
+
Requires-Dist: pandas>=1.3.0
|
|
18
|
+
Dynamic: author
|
|
19
|
+
Dynamic: license-file
|
|
20
|
+
Dynamic: requires-python
|
|
21
|
+
|
|
22
|
+
# BinBoost
|
|
23
|
+
|
|
24
|
+
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner, dengan Regularisasi Newton-Hessian dan Dukungan Multi-Kelas**
|
|
25
|
+
|
|
26
|
+
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR, termasuk negasi NOT) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Bobot tiap aturan dihitung dengan pendekatan Newton (memakai Hessian per fungsi loss) dan regularisasi adaptif berbasis dukungan sampel & kompleksitas aturan. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
27
|
+
|
|
28
|
+
## Kebaruan Utama
|
|
29
|
+
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
30
|
+
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti: `(A AND B)`, `(C OR D)`, atau `(A AND NOT B)`
|
|
31
|
+
- **Youden J**: Penggunaan indeks Youden J sebagai `threshold='auto'` bawaan untuk menentukan titik potong prediksi, bukan fixed di 0.5 serta mengoptimalkan sensitivity + specificity pada data training, cocok untuk kasus kelas tidak seimbang
|
|
32
|
+
- **Fitur ternegasi (NOT)**: pencarian aturan otomatis mempertimbangkan dua polaritas (asli dan negasi) untuk tiap kandidat fitur
|
|
33
|
+
- **Bobot aturan berbasis Newton (Hessian)**: bobot dua sisi tiap aturan dihitung dengan `w = G/(H+λ)`, memakai turunan kedua (Hessian) yang sesuai dengan fungsi loss yang dipilih (logistic, focal, atau poly), bukan sekadar rata-rata gradien
|
|
34
|
+
- **Regularisasi adaptif**: `λ` menyesuaikan otomatis terhadap jumlah sampel dan panjang aturan, sehingga aturan dengan cakupan kecil/kompleks ditahan bobotnya secara proporsional
|
|
35
|
+
- **Minimum-gain pruning (`gamma`)**: aturan dengan kontribusi terlalu kecil bisa ditolak sebelum masuk ensemble, membantu mencegah overfitting pada aturan yang terlalu spesifik
|
|
36
|
+
- **Dukungan multi-kelas**: lewat `binboost.multiclass.OneVsRestBinBoost`, BinBoost bisa dipakai untuk klasifikasi dengan lebih dari 2 kelas (strategi One-vs-Rest)
|
|
37
|
+
|
|
38
|
+
## Instalasi
|
|
39
|
+
```bash
|
|
40
|
+
pip install binboost
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Penggunaan Dasar untuk Klasifikasi Biner
|
|
44
|
+
```python
|
|
45
|
+
import numpy as np
|
|
46
|
+
from binboost import BinBoost
|
|
47
|
+
|
|
48
|
+
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
49
|
+
y = np.array([1, 0, 1, 0])
|
|
50
|
+
|
|
51
|
+
model = BinBoost(n_estimators=100, learning_rate=0.2, max_rule_length=4)
|
|
52
|
+
model.fit(X, y)
|
|
53
|
+
|
|
54
|
+
print(model.predict(X))
|
|
55
|
+
print(model.predict_proba(X))
|
|
56
|
+
print(model.rules_)
|
|
57
|
+
|
|
58
|
+
import pandas as pd
|
|
59
|
+
print(pd.DataFrame(model.rule_summary_))
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
## Penggunaan Multi-Kelas (>2 kelas)
|
|
63
|
+
Untuk label dengan lebih dari 2 kelas yang saling eksklusif, gunakan `OneVsRestBinBoost` — hyperparameter yang diteruskan persis sama dengan `BinBoost`:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
from binboost.multiclass import OneVsRestBinBoost
|
|
67
|
+
|
|
68
|
+
model = OneVsRestBinBoost(
|
|
69
|
+
mode='multiclass',
|
|
70
|
+
n_estimators=100,
|
|
71
|
+
learning_rate=0.2,
|
|
72
|
+
max_rule_length=4,
|
|
73
|
+
lambda0=3.0,
|
|
74
|
+
)
|
|
75
|
+
model.fit(X, y) # y boleh berisi >2 kelas, mis. array string atau integer
|
|
76
|
+
|
|
77
|
+
print(model.predict(X))
|
|
78
|
+
print(model.predict_proba(X)) # dinormalisasi supaya tiap baris berjumlah 1
|
|
79
|
+
print(pd.DataFrame(model.rule_summary_)) # aturan per kelas, kolom 'kelas' menandai sub-model asalnya
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
`mode='multilabel'` tersedia untuk kasus di mana satu sampel bisa punya lebih dari satu label positif sekaligus dimana probabilitas tidak dinormalisasi, tiap kelas diputuskan independen.
|
|
83
|
+
|
|
84
|
+
## Hyperparameter Utama
|
|
85
|
+
|
|
86
|
+
| Parameter | Bawaan | Keterangan |
|
|
87
|
+
|---|---|---|
|
|
88
|
+
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
89
|
+
| `learning_rate` | 0.2 | Faktor penyusutan tiap aturan |
|
|
90
|
+
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
91
|
+
| `max_rule_length` | 4 | Jumlah maksimum fitur dalam satu aturan |
|
|
92
|
+
| `operators` | `['AND','OR']` | Operator logika yang digunakan ubah `use_xor=True` untuk mengaktifkan XOR |
|
|
93
|
+
| `beam_width` | 5 | Lebar beam search |
|
|
94
|
+
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
95
|
+
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
96
|
+
| `feature_selection_threshold` | 0.0 | Ambang batas persentil gain fitur untuk seleksi kandidat (0.0 = semua fitur) |
|
|
97
|
+
| `min_samples_rule` | 5 | Jumlah/fraksi minimum sampel yang harus memenuhi sebuah aturan |
|
|
98
|
+
| `lambda0` | 3.0 | Konstanta regularisasi Newton adaptif |
|
|
99
|
+
| `gamma` | 0.0 | Ambang minimum gain untuk menerima sebuah aturan (0.0 = tanpa pruning, nonaktif secara bawaan) |
|
|
100
|
+
|
|
101
|
+
Fitur negasi (NOT) aktif otomatis pada pencarian aturan dan bukan hyperparameter yang bisa diatur lewat `BinBoost()` — ini bagian tetap dari mekanisme beam search.
|
|
102
|
+
|
|
103
|
+
## Fitur yang Didukung
|
|
104
|
+
- Fitur biner (0/1): langsung diproses
|
|
105
|
+
- Fitur numerik (int/float): dibinarisasi otomatis per iterasi
|
|
106
|
+
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
107
|
+
|
|
108
|
+
```python
|
|
109
|
+
from sklearn.preprocessing import OneHotEncoder
|
|
110
|
+
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
111
|
+
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
115
|
+
|
|
116
|
+
## Atribut Setelah Pelatihan
|
|
117
|
+
```python
|
|
118
|
+
model.rules_ # daftar teks aturan termasuk NOT jika ada
|
|
119
|
+
model.rule_weights_ # bobot dua sisi [w0, w1] tiap aturan (hasil Newton + regularisasi)
|
|
120
|
+
model.feature_importances_ # skor kepentingan fitur
|
|
121
|
+
model.train_score_ # loss per iterasi
|
|
122
|
+
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
123
|
+
model.n_rules_ # jumlah aturan aktif
|
|
124
|
+
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
125
|
+
model.thresholds_ # riwayat threshold biner per fitur numerik, per iterasi
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
## Lisensi
|
|
129
|
+
MIT
|
binboost-0.2.3/README.md
ADDED
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# BinBoost
|
|
2
|
+
|
|
3
|
+
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner, dengan Regularisasi Newton-Hessian dan Dukungan Multi-Kelas**
|
|
4
|
+
|
|
5
|
+
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR, termasuk negasi NOT) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Bobot tiap aturan dihitung dengan pendekatan Newton (memakai Hessian per fungsi loss) dan regularisasi adaptif berbasis dukungan sampel & kompleksitas aturan. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
6
|
+
|
|
7
|
+
## Kebaruan Utama
|
|
8
|
+
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
9
|
+
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti: `(A AND B)`, `(C OR D)`, atau `(A AND NOT B)`
|
|
10
|
+
- **Youden J**: Penggunaan indeks Youden J sebagai `threshold='auto'` bawaan untuk menentukan titik potong prediksi, bukan fixed di 0.5 serta mengoptimalkan sensitivity + specificity pada data training, cocok untuk kasus kelas tidak seimbang
|
|
11
|
+
- **Fitur ternegasi (NOT)**: pencarian aturan otomatis mempertimbangkan dua polaritas (asli dan negasi) untuk tiap kandidat fitur
|
|
12
|
+
- **Bobot aturan berbasis Newton (Hessian)**: bobot dua sisi tiap aturan dihitung dengan `w = G/(H+λ)`, memakai turunan kedua (Hessian) yang sesuai dengan fungsi loss yang dipilih (logistic, focal, atau poly), bukan sekadar rata-rata gradien
|
|
13
|
+
- **Regularisasi adaptif**: `λ` menyesuaikan otomatis terhadap jumlah sampel dan panjang aturan, sehingga aturan dengan cakupan kecil/kompleks ditahan bobotnya secara proporsional
|
|
14
|
+
- **Minimum-gain pruning (`gamma`)**: aturan dengan kontribusi terlalu kecil bisa ditolak sebelum masuk ensemble, membantu mencegah overfitting pada aturan yang terlalu spesifik
|
|
15
|
+
- **Dukungan multi-kelas**: lewat `binboost.multiclass.OneVsRestBinBoost`, BinBoost bisa dipakai untuk klasifikasi dengan lebih dari 2 kelas (strategi One-vs-Rest)
|
|
16
|
+
|
|
17
|
+
## Instalasi
|
|
18
|
+
```bash
|
|
19
|
+
pip install binboost
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
## Penggunaan Dasar untuk Klasifikasi Biner
|
|
23
|
+
```python
|
|
24
|
+
import numpy as np
|
|
25
|
+
from binboost import BinBoost
|
|
26
|
+
|
|
27
|
+
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
28
|
+
y = np.array([1, 0, 1, 0])
|
|
29
|
+
|
|
30
|
+
model = BinBoost(n_estimators=100, learning_rate=0.2, max_rule_length=4)
|
|
31
|
+
model.fit(X, y)
|
|
32
|
+
|
|
33
|
+
print(model.predict(X))
|
|
34
|
+
print(model.predict_proba(X))
|
|
35
|
+
print(model.rules_)
|
|
36
|
+
|
|
37
|
+
import pandas as pd
|
|
38
|
+
print(pd.DataFrame(model.rule_summary_))
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
## Penggunaan Multi-Kelas (>2 kelas)
|
|
42
|
+
Untuk label dengan lebih dari 2 kelas yang saling eksklusif, gunakan `OneVsRestBinBoost` — hyperparameter yang diteruskan persis sama dengan `BinBoost`:
|
|
43
|
+
|
|
44
|
+
```python
|
|
45
|
+
from binboost.multiclass import OneVsRestBinBoost
|
|
46
|
+
|
|
47
|
+
model = OneVsRestBinBoost(
|
|
48
|
+
mode='multiclass',
|
|
49
|
+
n_estimators=100,
|
|
50
|
+
learning_rate=0.2,
|
|
51
|
+
max_rule_length=4,
|
|
52
|
+
lambda0=3.0,
|
|
53
|
+
)
|
|
54
|
+
model.fit(X, y) # y boleh berisi >2 kelas, mis. array string atau integer
|
|
55
|
+
|
|
56
|
+
print(model.predict(X))
|
|
57
|
+
print(model.predict_proba(X)) # dinormalisasi supaya tiap baris berjumlah 1
|
|
58
|
+
print(pd.DataFrame(model.rule_summary_)) # aturan per kelas, kolom 'kelas' menandai sub-model asalnya
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
`mode='multilabel'` tersedia untuk kasus di mana satu sampel bisa punya lebih dari satu label positif sekaligus dimana probabilitas tidak dinormalisasi, tiap kelas diputuskan independen.
|
|
62
|
+
|
|
63
|
+
## Hyperparameter Utama
|
|
64
|
+
|
|
65
|
+
| Parameter | Bawaan | Keterangan |
|
|
66
|
+
|---|---|---|
|
|
67
|
+
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
68
|
+
| `learning_rate` | 0.2 | Faktor penyusutan tiap aturan |
|
|
69
|
+
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
70
|
+
| `max_rule_length` | 4 | Jumlah maksimum fitur dalam satu aturan |
|
|
71
|
+
| `operators` | `['AND','OR']` | Operator logika yang digunakan ubah `use_xor=True` untuk mengaktifkan XOR |
|
|
72
|
+
| `beam_width` | 5 | Lebar beam search |
|
|
73
|
+
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
74
|
+
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
75
|
+
| `feature_selection_threshold` | 0.0 | Ambang batas persentil gain fitur untuk seleksi kandidat (0.0 = semua fitur) |
|
|
76
|
+
| `min_samples_rule` | 5 | Jumlah/fraksi minimum sampel yang harus memenuhi sebuah aturan |
|
|
77
|
+
| `lambda0` | 3.0 | Konstanta regularisasi Newton adaptif |
|
|
78
|
+
| `gamma` | 0.0 | Ambang minimum gain untuk menerima sebuah aturan (0.0 = tanpa pruning, nonaktif secara bawaan) |
|
|
79
|
+
|
|
80
|
+
Fitur negasi (NOT) aktif otomatis pada pencarian aturan dan bukan hyperparameter yang bisa diatur lewat `BinBoost()` — ini bagian tetap dari mekanisme beam search.
|
|
81
|
+
|
|
82
|
+
## Fitur yang Didukung
|
|
83
|
+
- Fitur biner (0/1): langsung diproses
|
|
84
|
+
- Fitur numerik (int/float): dibinarisasi otomatis per iterasi
|
|
85
|
+
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
86
|
+
|
|
87
|
+
```python
|
|
88
|
+
from sklearn.preprocessing import OneHotEncoder
|
|
89
|
+
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
90
|
+
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
94
|
+
|
|
95
|
+
## Atribut Setelah Pelatihan
|
|
96
|
+
```python
|
|
97
|
+
model.rules_ # daftar teks aturan termasuk NOT jika ada
|
|
98
|
+
model.rule_weights_ # bobot dua sisi [w0, w1] tiap aturan (hasil Newton + regularisasi)
|
|
99
|
+
model.feature_importances_ # skor kepentingan fitur
|
|
100
|
+
model.train_score_ # loss per iterasi
|
|
101
|
+
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
102
|
+
model.n_rules_ # jumlah aturan aktif
|
|
103
|
+
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
104
|
+
model.thresholds_ # riwayat threshold biner per fitur numerik, per iterasi
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
## Lisensi
|
|
108
|
+
MIT
|
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
# binboost/binboost.py
|
|
1
2
|
import numpy as np
|
|
2
3
|
|
|
3
4
|
from .loss import get_loss, sigmoid
|
|
@@ -9,7 +10,7 @@ class BinBoost:
|
|
|
9
10
|
"""
|
|
10
11
|
BinBoost: Gradient Boosting Berbasis Aturan Logika Adaptif.
|
|
11
12
|
|
|
12
|
-
Algoritma klasifikasi biner yang membangun ensemble aturan logika
|
|
13
|
+
Algoritma klasifikasi biner yang membangun ensemble aturan logika
|
|
13
14
|
murni (AND, OR, XOR) dengan binarisasi fitur numerik adaptif
|
|
14
15
|
berbasis gradien pada setiap iterasi boosting.
|
|
15
16
|
|
|
@@ -18,7 +19,7 @@ class BinBoost:
|
|
|
18
19
|
n_estimators : int, default=100
|
|
19
20
|
Jumlah iterasi boosting.
|
|
20
21
|
|
|
21
|
-
learning_rate : float, default=0.
|
|
22
|
+
learning_rate : float, default=0.2
|
|
22
23
|
Faktor penyusutan kontribusi setiap aturan.
|
|
23
24
|
|
|
24
25
|
loss : str, default='logistic'
|
|
@@ -33,7 +34,7 @@ class BinBoost:
|
|
|
33
34
|
focal_alpha : float, default=0.25
|
|
34
35
|
Parameter alpha pada Focal Loss. Hanya berlaku saat loss='focal'.
|
|
35
36
|
|
|
36
|
-
max_rule_length : int, default=
|
|
37
|
+
max_rule_length : int, default=4
|
|
37
38
|
Jumlah maksimum fitur dalam satu aturan.
|
|
38
39
|
|
|
39
40
|
operators : list, default=['AND', 'OR']
|
|
@@ -59,7 +60,7 @@ class BinBoost:
|
|
|
59
60
|
min_samples_rule : int or float, default=5
|
|
60
61
|
Jumlah minimum sampel yang harus memenuhi sebuah aturan.
|
|
61
62
|
|
|
62
|
-
lambda0 : float, default=
|
|
63
|
+
lambda0 : float, default=3.0
|
|
63
64
|
Konstanta regularisasi dasar untuk bobot Newton adaptif (λ_R = lambda0 · L / sqrt(n_R+1)).
|
|
64
65
|
|
|
65
66
|
binarize_strategy : str, default='gradient'
|
|
@@ -122,12 +123,12 @@ class BinBoost:
|
|
|
122
123
|
def __init__(
|
|
123
124
|
self,
|
|
124
125
|
n_estimators=100,
|
|
125
|
-
learning_rate=0.
|
|
126
|
+
learning_rate=0.2,
|
|
126
127
|
loss='logistic',
|
|
127
128
|
poly_epsilon=1.0,
|
|
128
129
|
focal_gamma=2.0,
|
|
129
130
|
focal_alpha=0.25,
|
|
130
|
-
max_rule_length=
|
|
131
|
+
max_rule_length=4,
|
|
131
132
|
operators=None,
|
|
132
133
|
use_xor=False,
|
|
133
134
|
rule_complexity_penalty=0.0,
|
|
@@ -135,7 +136,8 @@ class BinBoost:
|
|
|
135
136
|
beam_width=5,
|
|
136
137
|
subsample=0.8,
|
|
137
138
|
min_samples_rule=5,
|
|
138
|
-
lambda0=
|
|
139
|
+
lambda0=3.0,
|
|
140
|
+
gamma=0.0,
|
|
139
141
|
binarize_strategy='gradient',
|
|
140
142
|
n_thresholds='auto',
|
|
141
143
|
n_iter_no_change=None,
|
|
@@ -159,6 +161,7 @@ class BinBoost:
|
|
|
159
161
|
self.subsample = subsample
|
|
160
162
|
self.min_samples_rule = min_samples_rule
|
|
161
163
|
self.lambda0 = lambda0
|
|
164
|
+
self.gamma = gamma
|
|
162
165
|
self.binarize_strategy = binarize_strategy
|
|
163
166
|
self.n_thresholds = n_thresholds
|
|
164
167
|
self.n_iter_no_change = n_iter_no_change
|
|
@@ -322,6 +325,7 @@ class BinBoost:
|
|
|
322
325
|
min_samples_rule=self.min_samples_rule,
|
|
323
326
|
rule_complexity_penalty=self.rule_complexity_penalty,
|
|
324
327
|
ohe_groups=self.ohe_groups_,
|
|
328
|
+
gamma=self.gamma,
|
|
325
329
|
)
|
|
326
330
|
|
|
327
331
|
if not self.warm_start or not hasattr(self, 'estimators_'):
|
|
@@ -438,16 +442,25 @@ class BinBoost:
|
|
|
438
442
|
skor_terbaik = youden
|
|
439
443
|
threshold_terbaik = t
|
|
440
444
|
return float(threshold_terbaik)
|
|
445
|
+
|
|
446
|
+
def _binarize_untuk_iterasi(self, X, iterasi_idx, numeric_cols):
|
|
447
|
+
X_bin = X.copy().astype(np.float64)
|
|
448
|
+
for col in numeric_cols:
|
|
449
|
+
t = self.thresholds_[col][iterasi_idx]
|
|
450
|
+
X_bin[:, col] = (X[:, col] > t).astype(np.float64)
|
|
451
|
+
return X_bin
|
|
441
452
|
|
|
442
453
|
def _decision_function(self, X):
|
|
443
|
-
# Hitung nilai F akhir untuk seluruh sampel menggunakan update dua sisi
|
|
444
454
|
self._check_is_fitted()
|
|
445
455
|
X, _ = self._validate_input(X)
|
|
446
456
|
F = np.zeros(X.shape[0])
|
|
447
457
|
numeric_cols = self._detect_numeric_cols(X)
|
|
448
|
-
|
|
449
|
-
|
|
450
|
-
|
|
458
|
+
for i, (rule, weights) in enumerate(zip(self.estimators_, self.rule_weights_)):
|
|
459
|
+
if self.binarize_strategy == 'gradient' and numeric_cols:
|
|
460
|
+
X_bin = self._binarize_untuk_iterasi(X, i, numeric_cols)
|
|
461
|
+
else:
|
|
462
|
+
dummy_grad = np.ones(X.shape[0])
|
|
463
|
+
X_bin, _ = self._binarizer.transform(X, dummy_grad, numeric_cols)
|
|
451
464
|
h = rule.evaluate(X_bin)
|
|
452
465
|
w0, w1 = weights[0], weights[1]
|
|
453
466
|
F += self.learning_rate * (w1 * h + w0 * (1.0 - h))
|
|
@@ -509,10 +522,13 @@ class BinBoost:
|
|
|
509
522
|
self._check_is_fitted()
|
|
510
523
|
X, _ = self._validate_input(X)
|
|
511
524
|
numeric_cols = self._detect_numeric_cols(X)
|
|
512
|
-
dummy_grad = np.ones(X.shape[0])
|
|
513
|
-
X_bin, _ = self._binarizer.transform(X, dummy_grad, numeric_cols)
|
|
514
525
|
F = np.zeros(X.shape[0])
|
|
515
|
-
for rule, weights in zip(self.estimators_, self.rule_weights_):
|
|
526
|
+
for i, (rule, weights) in enumerate(zip(self.estimators_, self.rule_weights_)):
|
|
527
|
+
if self.binarize_strategy == 'gradient' and numeric_cols:
|
|
528
|
+
X_bin = self._binarize_untuk_iterasi(X, i, numeric_cols)
|
|
529
|
+
else:
|
|
530
|
+
dummy_grad = np.ones(X.shape[0])
|
|
531
|
+
X_bin, _ = self._binarizer.transform(X, dummy_grad, numeric_cols)
|
|
516
532
|
h = rule.evaluate(X_bin)
|
|
517
533
|
w0, w1 = weights[0], weights[1]
|
|
518
534
|
F += self.learning_rate * (w1 * h + w0 * (1.0 - h))
|
|
@@ -533,7 +549,7 @@ class BinBoost:
|
|
|
533
549
|
def apply(self, X):
|
|
534
550
|
"""
|
|
535
551
|
Kembalikan nilai aktivasi setiap aturan untuk setiap sampel.
|
|
536
|
-
|
|
552
|
+
|
|
537
553
|
Kembalian
|
|
538
554
|
----------
|
|
539
555
|
ndarray of shape (n_samples, n_estimators_)
|
|
@@ -541,11 +557,16 @@ class BinBoost:
|
|
|
541
557
|
self._check_is_fitted()
|
|
542
558
|
X, _ = self._validate_input(X)
|
|
543
559
|
numeric_cols = self._detect_numeric_cols(X)
|
|
544
|
-
dummy_grad = np.ones(X.shape[0])
|
|
545
|
-
X_bin, _ = self._binarizer.transform(X, dummy_grad, numeric_cols)
|
|
546
560
|
results = []
|
|
547
|
-
|
|
548
|
-
|
|
561
|
+
|
|
562
|
+
for i, rule in enumerate(self.estimators_):
|
|
563
|
+
if self.binarize_strategy == 'gradient' and numeric_cols:
|
|
564
|
+
X_bin = self._binarize_untuk_iterasi(X, i, numeric_cols)
|
|
565
|
+
else:
|
|
566
|
+
dummy_grad = np.ones(X.shape[0])
|
|
567
|
+
X_bin, _ = self._binarizer.transform(X, dummy_grad, numeric_cols)
|
|
568
|
+
results.append(rule.evaluate(X_bin))
|
|
569
|
+
|
|
549
570
|
return np.column_stack(results)
|
|
550
571
|
|
|
551
572
|
def get_params(self, deep=True):
|
|
@@ -571,7 +592,8 @@ class BinBoost:
|
|
|
571
592
|
'beam_width': self.beam_width,
|
|
572
593
|
'subsample': self.subsample,
|
|
573
594
|
'min_samples_rule': self.min_samples_rule,
|
|
574
|
-
'lambda0': self.lambda0,
|
|
595
|
+
'lambda0': self.lambda0,
|
|
596
|
+
'gamma': self.gamma,
|
|
575
597
|
'binarize_strategy': self.binarize_strategy,
|
|
576
598
|
'n_thresholds': self.n_thresholds,
|
|
577
599
|
'n_iter_no_change': self.n_iter_no_change,
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# binboost/multiclass.py
|
|
2
|
+
import numpy as np
|
|
3
|
+
from .binboost import BinBoost
|
|
4
|
+
|
|
5
|
+
|
|
6
|
+
class OneVsRestBinBoost:
|
|
7
|
+
"""
|
|
8
|
+
Ekstensi BinBoost untuk klasifikasi multi-kelas (>2 kelas) lewat strategi
|
|
9
|
+
One-vs-Rest, mirip pola sklearn.multiclass.OneVsRestClassifier.
|
|
10
|
+
Untuk 2 kelas, gunakan langsung BinBoost.
|
|
11
|
+
"""
|
|
12
|
+
|
|
13
|
+
def __init__(self, mode='multiclass', **binboost_params):
|
|
14
|
+
self.mode = mode
|
|
15
|
+
self.binboost_params = binboost_params
|
|
16
|
+
|
|
17
|
+
def fit(self, X, y):
|
|
18
|
+
X = np.asarray(X, dtype=np.float64)
|
|
19
|
+
y = np.asarray(y)
|
|
20
|
+
self.classes_ = np.unique(y)
|
|
21
|
+
self.n_classes_ = len(self.classes_)
|
|
22
|
+
if self.n_classes_ < 2:
|
|
23
|
+
raise ValueError("Minimal harus ada 2 kelas berbeda pada label y.")
|
|
24
|
+
self.estimators_ovr_ = []
|
|
25
|
+
for kelas in self.classes_:
|
|
26
|
+
y_bin = (y == kelas).astype(np.float64)
|
|
27
|
+
model_k = BinBoost(**self.binboost_params)
|
|
28
|
+
model_k.fit(X, y_bin)
|
|
29
|
+
self.estimators_ovr_.append(model_k)
|
|
30
|
+
return self
|
|
31
|
+
|
|
32
|
+
def _raw_scores(self, X):
|
|
33
|
+
return np.column_stack([m.predict_proba(X)[:, 1] for m in self.estimators_ovr_])
|
|
34
|
+
|
|
35
|
+
def predict_proba(self, X):
|
|
36
|
+
raw = self._raw_scores(X)
|
|
37
|
+
if self.mode == 'multilabel':
|
|
38
|
+
return raw
|
|
39
|
+
total = raw.sum(axis=1, keepdims=True)
|
|
40
|
+
semua_nol = (total.flatten() <= 1e-10)
|
|
41
|
+
total_aman = np.where(total <= 1e-10, 1.0, total)
|
|
42
|
+
probs = raw / total_aman
|
|
43
|
+
if semua_nol.any():
|
|
44
|
+
probs[semua_nol] = 1.0 / self.n_classes_
|
|
45
|
+
return probs
|
|
46
|
+
|
|
47
|
+
def predict(self, X):
|
|
48
|
+
if self.mode == 'multilabel':
|
|
49
|
+
return np.column_stack([m.predict(X) for m in self.estimators_ovr_])
|
|
50
|
+
probs = self.predict_proba(X)
|
|
51
|
+
idx = np.argmax(probs, axis=1)
|
|
52
|
+
return self.classes_[idx]
|
|
53
|
+
|
|
54
|
+
def score(self, X, y):
|
|
55
|
+
y = np.asarray(y)
|
|
56
|
+
return np.mean(self.predict(X) == y)
|
|
57
|
+
|
|
58
|
+
@property
|
|
59
|
+
def rule_summary_(self):
|
|
60
|
+
rows = []
|
|
61
|
+
for kelas, model_k in zip(self.classes_, self.estimators_ovr_):
|
|
62
|
+
for r in model_k.rule_summary_:
|
|
63
|
+
r = dict(r)
|
|
64
|
+
r['kelas'] = kelas
|
|
65
|
+
rows.append(r)
|
|
66
|
+
return rows
|
|
67
|
+
|
|
68
|
+
@property
|
|
69
|
+
def feature_usage_(self):
|
|
70
|
+
usage_total = {}
|
|
71
|
+
for model_k in self.estimators_ovr_:
|
|
72
|
+
for k, v in model_k.feature_usage_.items():
|
|
73
|
+
usage_total[k] = usage_total.get(k, 0) + v
|
|
74
|
+
return usage_total
|
|
75
|
+
|
|
76
|
+
@property
|
|
77
|
+
def n_rules_(self):
|
|
78
|
+
return sum(m.n_rules_ for m in self.estimators_ovr_)
|
|
@@ -1,23 +1,26 @@
|
|
|
1
|
-
# rule.py
|
|
1
|
+
# binboost/rule.py
|
|
2
2
|
import numpy as np
|
|
3
3
|
|
|
4
4
|
class Rule:
|
|
5
|
-
# Satu aturan logika dalam ensemble BinBoost
|
|
6
|
-
def __init__(self, features, operators, weight, feature_names=None):
|
|
7
|
-
# 1. Simpan komponen aturan
|
|
8
|
-
# a. Indeks fitur yang terlibat dalam aturan
|
|
5
|
+
# Satu aturan logika dalam ensemble BinBoost, dengan dukungan negasi (NOT) per fitur
|
|
6
|
+
def __init__(self, features, operators, weight, negations=None, feature_names=None):
|
|
9
7
|
self.features = features
|
|
10
|
-
# b. Daftar operator logika antar fitur
|
|
11
8
|
self.operators = operators
|
|
12
|
-
|
|
13
|
-
self.
|
|
9
|
+
self._search_weight = weight
|
|
10
|
+
self.negations = negations if negations is not None else [False] * len(features)
|
|
14
11
|
self.feature_names = feature_names
|
|
15
12
|
|
|
13
|
+
def _kolom(self, X_bin, idx_dalam_features):
|
|
14
|
+
fitur_idx = self.features[idx_dalam_features]
|
|
15
|
+
kolom = X_bin[:, fitur_idx].astype(bool)
|
|
16
|
+
if self.negations[idx_dalam_features]:
|
|
17
|
+
kolom = ~kolom
|
|
18
|
+
return kolom
|
|
19
|
+
|
|
16
20
|
def evaluate(self, X_bin):
|
|
17
|
-
|
|
18
|
-
result = X_bin[:, self.features[0]].astype(bool)
|
|
21
|
+
result = self._kolom(X_bin, 0)
|
|
19
22
|
for i, op in enumerate(self.operators):
|
|
20
|
-
next_col = X_bin
|
|
23
|
+
next_col = self._kolom(X_bin, i + 1)
|
|
21
24
|
if op == 'AND':
|
|
22
25
|
result = result & next_col
|
|
23
26
|
elif op == 'OR':
|
|
@@ -27,12 +30,12 @@ class Rule:
|
|
|
27
30
|
return result.astype(np.float64)
|
|
28
31
|
|
|
29
32
|
def to_string(self, feature_names=None):
|
|
30
|
-
# Kembalikan representasi teks aturan yang dapat dibaca manusia
|
|
31
33
|
names = feature_names if feature_names is not None else self.feature_names
|
|
32
34
|
if names is None:
|
|
33
|
-
|
|
35
|
+
base_parts = [f"X{f}" for f in self.features]
|
|
34
36
|
else:
|
|
35
|
-
|
|
37
|
+
base_parts = [str(names[f]) for f in self.features]
|
|
38
|
+
parts = [f"NOT {p}" if neg else p for p, neg in zip(base_parts, self.negations)]
|
|
36
39
|
if len(parts) == 1:
|
|
37
40
|
return parts[0]
|
|
38
41
|
expr = parts[0]
|
|
@@ -41,15 +44,13 @@ class Rule:
|
|
|
41
44
|
return expr
|
|
42
45
|
|
|
43
46
|
def __repr__(self):
|
|
44
|
-
return f"Rule({self.to_string()}
|
|
47
|
+
return f"Rule({self.to_string()})"
|
|
45
48
|
|
|
46
49
|
|
|
47
50
|
class BeamSearchRuleFinder:
|
|
48
|
-
# Pencari aturan terbaik menggunakan beam search tervektorisasi berbasis numpy
|
|
49
|
-
|
|
50
51
|
def __init__(self, operators, max_rule_length, beam_width,
|
|
51
52
|
feature_selection_threshold, min_samples_rule,
|
|
52
|
-
rule_complexity_penalty, ohe_groups=None):
|
|
53
|
+
rule_complexity_penalty, ohe_groups=None, allow_negation=True, gamma=0.0):
|
|
53
54
|
self.operators = operators
|
|
54
55
|
self.max_rule_length = max_rule_length
|
|
55
56
|
self.beam_width = beam_width
|
|
@@ -57,8 +58,9 @@ class BeamSearchRuleFinder:
|
|
|
57
58
|
self.min_samples_rule = min_samples_rule
|
|
58
59
|
self.rule_complexity_penalty = rule_complexity_penalty
|
|
59
60
|
self.ohe_groups = ohe_groups
|
|
61
|
+
self.allow_negation = allow_negation
|
|
62
|
+
self.gamma = gamma
|
|
60
63
|
|
|
61
|
-
# Bangun peta indeks fitur ke grup OHE untuk pencarian constraint yang cepat
|
|
62
64
|
self._feature_to_group = {}
|
|
63
65
|
if ohe_groups is not None:
|
|
64
66
|
for gid, members in ohe_groups.items():
|
|
@@ -66,8 +68,6 @@ class BeamSearchRuleFinder:
|
|
|
66
68
|
self._feature_to_group[f] = gid
|
|
67
69
|
|
|
68
70
|
def _select_features(self, X_bin, gradients):
|
|
69
|
-
# 1. Pilih fitur kandidat berdasarkan skor korelasi gradien secara tervektorisasi
|
|
70
|
-
# a. Hitung dot product gradien dengan semua kolom sekaligus dalam satu operasi
|
|
71
71
|
dot_products = np.abs(X_bin.T @ gradients)
|
|
72
72
|
col_sums = X_bin.sum(axis=0) + 1e-10
|
|
73
73
|
scores = dot_products / col_sums
|
|
@@ -77,29 +77,24 @@ class BeamSearchRuleFinder:
|
|
|
77
77
|
return selected
|
|
78
78
|
|
|
79
79
|
def _compute_scores_batch(self, H_batch, gradients):
|
|
80
|
-
# 1. Hitung bobot optimal dan skor untuk banyak kandidat aturan sekaligus
|
|
81
|
-
# a. H_batch berukuran (n_kandidat, n_sampel)
|
|
82
80
|
dot_gh = H_batch @ gradients
|
|
83
81
|
dot_hh = (H_batch * H_batch).sum(axis=1)
|
|
84
82
|
valid = dot_hh > 1e-10
|
|
85
83
|
w = np.where(valid, dot_gh / np.where(dot_hh > 1e-10, dot_hh, 1.0), 0.0)
|
|
86
|
-
# b. Kurangi gradien dengan prediksi berbobot untuk semua kandidat sekaligus
|
|
87
84
|
residual = gradients[np.newaxis, :] - w[:, np.newaxis] * H_batch
|
|
88
85
|
scores = (residual * residual).sum(axis=1)
|
|
89
86
|
scores = np.where(valid, scores, np.inf)
|
|
90
87
|
return w, scores, valid
|
|
91
88
|
|
|
92
89
|
def _min_samples_int(self, n_samples):
|
|
93
|
-
# Kembalikan jumlah sampel minimum dalam bentuk bilangan bulat
|
|
94
90
|
if isinstance(self.min_samples_rule, float) and self.min_samples_rule < 1.0:
|
|
95
91
|
return max(1, int(self.min_samples_rule * n_samples))
|
|
96
92
|
return int(self.min_samples_rule)
|
|
97
93
|
|
|
98
94
|
def find_best_rule(self, X_bin, gradients):
|
|
99
|
-
# 1. Jalankan beam search tervektorisasi untuk menemukan aturan terbaik
|
|
100
95
|
n_samples, n_features = X_bin.shape
|
|
101
96
|
min_s = self._min_samples_int(n_samples)
|
|
102
|
-
|
|
97
|
+
baseline = float(np.dot(gradients, gradients))
|
|
103
98
|
if hasattr(self, 'forced_candidate_features') and self.forced_candidate_features is not None:
|
|
104
99
|
candidate_features = self.forced_candidate_features
|
|
105
100
|
else:
|
|
@@ -109,42 +104,49 @@ class BeamSearchRuleFinder:
|
|
|
109
104
|
return None
|
|
110
105
|
|
|
111
106
|
X_bin_bool = X_bin.astype(bool)
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
107
|
+
polaritas_list = [False, True] if self.allow_negation else [False]
|
|
108
|
+
|
|
109
|
+
# 1. Inisialisasi beam panjang satu — coba dua polaritas (asli & negasi) per fitur
|
|
110
|
+
H_list, meta_list = [], []
|
|
111
|
+
for cf in candidate_features:
|
|
112
|
+
kolom_pos = X_bin_bool[:, int(cf)]
|
|
113
|
+
for neg in polaritas_list:
|
|
114
|
+
kolom = ~kolom_pos if neg else kolom_pos
|
|
115
|
+
H_list.append(kolom.astype(np.float64))
|
|
116
|
+
meta_list.append((int(cf), neg))
|
|
117
|
+
|
|
118
|
+
H_init = np.array(H_list)
|
|
115
119
|
support_init = H_init.sum(axis=1)
|
|
116
120
|
valid_init = support_init >= min_s
|
|
117
|
-
|
|
118
121
|
if not valid_init.any():
|
|
119
122
|
return None
|
|
120
123
|
|
|
121
124
|
w_init, scores_init, w_valid = self._compute_scores_batch(H_init, gradients)
|
|
122
|
-
|
|
123
|
-
|
|
125
|
+
gain_init = baseline - scores_init
|
|
126
|
+
valid_mask = valid_init & w_valid & (gain_init >= self.gamma)
|
|
124
127
|
if not valid_mask.any():
|
|
125
128
|
return None
|
|
126
129
|
|
|
127
130
|
candidates = []
|
|
128
|
-
for i, cf in enumerate(
|
|
131
|
+
for i, (cf, neg) in enumerate(meta_list):
|
|
129
132
|
if valid_mask[i]:
|
|
130
133
|
pen = scores_init[i] * (1 + self.rule_complexity_penalty * 1)
|
|
131
|
-
candidates.append((pen, [
|
|
134
|
+
candidates.append((pen, [cf], [], [neg], H_init[i].astype(bool)))
|
|
132
135
|
|
|
133
136
|
if not candidates:
|
|
134
137
|
return None
|
|
135
138
|
|
|
136
139
|
candidates.sort(key=lambda x: x[0])
|
|
137
140
|
beam = candidates[:self.beam_width]
|
|
138
|
-
best_score, best_feats, best_ops, best_h = beam[0]
|
|
141
|
+
best_score, best_feats, best_ops, best_negs, best_h = beam[0]
|
|
139
142
|
|
|
140
|
-
#
|
|
143
|
+
# 2. Perluas aturan hingga panjang maksimum
|
|
141
144
|
for length in range(2, self.max_rule_length + 1):
|
|
142
145
|
new_candidates = []
|
|
143
146
|
|
|
144
|
-
for beam_score, beam_feats, beam_ops, beam_h in beam:
|
|
147
|
+
for beam_score, beam_feats, beam_ops, beam_negs, beam_h in beam:
|
|
145
148
|
beam_feat_set = set(beam_feats)
|
|
146
149
|
|
|
147
|
-
# a. Tentukan fitur yang dapat diperluas dengan mempertimbangkan constraint OHE
|
|
148
150
|
expandable = []
|
|
149
151
|
for cf in candidate_features:
|
|
150
152
|
cf_int = int(cf)
|
|
@@ -159,42 +161,43 @@ class BeamSearchRuleFinder:
|
|
|
159
161
|
if not expandable:
|
|
160
162
|
continue
|
|
161
163
|
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
valid_sup = support_new >= min_s
|
|
179
|
-
|
|
180
|
-
if not valid_sup.any():
|
|
181
|
-
continue
|
|
182
|
-
|
|
183
|
-
w_new, scores_new, w_valid_new = self._compute_scores_batch(
|
|
184
|
-
H_new, gradients
|
|
185
|
-
)
|
|
186
|
-
valid_combined = valid_sup & w_valid_new
|
|
164
|
+
cols_pos = X_bin_bool[:, expandable].T
|
|
165
|
+
kandidat_kolom = [(cols_pos, False)]
|
|
166
|
+
if self.allow_negation:
|
|
167
|
+
kandidat_kolom.append((~cols_pos, True))
|
|
168
|
+
|
|
169
|
+
for cols, neg_flag in kandidat_kolom:
|
|
170
|
+
for op in self.operators:
|
|
171
|
+
beam_h_2d = np.broadcast_to(beam_h, cols.shape)
|
|
172
|
+
if op == 'AND':
|
|
173
|
+
H_new = (beam_h_2d & cols).astype(np.float64)
|
|
174
|
+
elif op == 'OR':
|
|
175
|
+
H_new = (beam_h_2d | cols).astype(np.float64)
|
|
176
|
+
elif op == 'XOR':
|
|
177
|
+
H_new = (beam_h_2d ^ cols).astype(np.float64)
|
|
178
|
+
else:
|
|
179
|
+
continue
|
|
187
180
|
|
|
188
|
-
|
|
189
|
-
|
|
181
|
+
support_new = H_new.sum(axis=1)
|
|
182
|
+
valid_sup = support_new >= min_s
|
|
183
|
+
if not valid_sup.any():
|
|
190
184
|
continue
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
185
|
+
|
|
186
|
+
w_new, scores_new, w_valid_new = self._compute_scores_batch(H_new, gradients)
|
|
187
|
+
gain_new = baseline - scores_new
|
|
188
|
+
valid_combined = valid_sup & w_valid_new & (gain_new >= self.gamma)
|
|
189
|
+
|
|
190
|
+
for j, cf_int in enumerate(expandable):
|
|
191
|
+
if not valid_combined[j]:
|
|
192
|
+
continue
|
|
193
|
+
pen = scores_new[j] * (1 + self.rule_complexity_penalty * length)
|
|
194
|
+
new_candidates.append((
|
|
195
|
+
pen,
|
|
196
|
+
beam_feats + [cf_int],
|
|
197
|
+
beam_ops + [op],
|
|
198
|
+
beam_negs + [neg_flag],
|
|
199
|
+
H_new[j].astype(bool)
|
|
200
|
+
))
|
|
198
201
|
|
|
199
202
|
if not new_candidates:
|
|
200
203
|
break
|
|
@@ -203,16 +206,15 @@ class BeamSearchRuleFinder:
|
|
|
203
206
|
new_candidates = new_candidates[:self.beam_width]
|
|
204
207
|
|
|
205
208
|
if new_candidates[0][0] < best_score:
|
|
206
|
-
best_score, best_feats, best_ops, best_h = new_candidates[0]
|
|
209
|
+
best_score, best_feats, best_ops, best_negs, best_h = new_candidates[0]
|
|
207
210
|
beam = new_candidates
|
|
208
211
|
else:
|
|
209
212
|
break
|
|
210
213
|
|
|
211
|
-
# 4. Hitung ulang bobot optimal untuk aturan terbaik yang dipilih
|
|
212
214
|
h_final = best_h.astype(np.float64)
|
|
213
215
|
denom = np.dot(h_final, h_final)
|
|
214
216
|
if denom < 1e-10:
|
|
215
217
|
return None
|
|
216
218
|
w_final = np.dot(gradients, h_final) / denom
|
|
217
219
|
|
|
218
|
-
return Rule(best_feats, best_ops, w_final)
|
|
220
|
+
return Rule(best_feats, best_ops, w_final, negations=best_negs)
|
|
@@ -0,0 +1,129 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: binboost
|
|
3
|
+
Version: 0.2.3
|
|
4
|
+
Summary: Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner
|
|
5
|
+
Author: Rangga Wahyu Pratama
|
|
6
|
+
License: MIT
|
|
7
|
+
Keywords: gradient boosting,logical rules,binary features,interpretable machine learning
|
|
8
|
+
Classifier: Programming Language :: Python :: 3
|
|
9
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
+
Classifier: Operating System :: OS Independent
|
|
11
|
+
Classifier: Intended Audience :: Science/Research
|
|
12
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
13
|
+
Requires-Python: >=3.8
|
|
14
|
+
Description-Content-Type: text/markdown
|
|
15
|
+
License-File: LICENSE
|
|
16
|
+
Requires-Dist: numpy>=1.21.0
|
|
17
|
+
Requires-Dist: pandas>=1.3.0
|
|
18
|
+
Dynamic: author
|
|
19
|
+
Dynamic: license-file
|
|
20
|
+
Dynamic: requires-python
|
|
21
|
+
|
|
22
|
+
# BinBoost
|
|
23
|
+
|
|
24
|
+
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner, dengan Regularisasi Newton-Hessian dan Dukungan Multi-Kelas**
|
|
25
|
+
|
|
26
|
+
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR, termasuk negasi NOT) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Bobot tiap aturan dihitung dengan pendekatan Newton (memakai Hessian per fungsi loss) dan regularisasi adaptif berbasis dukungan sampel & kompleksitas aturan. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
27
|
+
|
|
28
|
+
## Kebaruan Utama
|
|
29
|
+
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
30
|
+
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti: `(A AND B)`, `(C OR D)`, atau `(A AND NOT B)`
|
|
31
|
+
- **Youden J**: Penggunaan indeks Youden J sebagai `threshold='auto'` bawaan untuk menentukan titik potong prediksi, bukan fixed di 0.5 serta mengoptimalkan sensitivity + specificity pada data training, cocok untuk kasus kelas tidak seimbang
|
|
32
|
+
- **Fitur ternegasi (NOT)**: pencarian aturan otomatis mempertimbangkan dua polaritas (asli dan negasi) untuk tiap kandidat fitur
|
|
33
|
+
- **Bobot aturan berbasis Newton (Hessian)**: bobot dua sisi tiap aturan dihitung dengan `w = G/(H+λ)`, memakai turunan kedua (Hessian) yang sesuai dengan fungsi loss yang dipilih (logistic, focal, atau poly), bukan sekadar rata-rata gradien
|
|
34
|
+
- **Regularisasi adaptif**: `λ` menyesuaikan otomatis terhadap jumlah sampel dan panjang aturan, sehingga aturan dengan cakupan kecil/kompleks ditahan bobotnya secara proporsional
|
|
35
|
+
- **Minimum-gain pruning (`gamma`)**: aturan dengan kontribusi terlalu kecil bisa ditolak sebelum masuk ensemble, membantu mencegah overfitting pada aturan yang terlalu spesifik
|
|
36
|
+
- **Dukungan multi-kelas**: lewat `binboost.multiclass.OneVsRestBinBoost`, BinBoost bisa dipakai untuk klasifikasi dengan lebih dari 2 kelas (strategi One-vs-Rest)
|
|
37
|
+
|
|
38
|
+
## Instalasi
|
|
39
|
+
```bash
|
|
40
|
+
pip install binboost
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## Penggunaan Dasar untuk Klasifikasi Biner
|
|
44
|
+
```python
|
|
45
|
+
import numpy as np
|
|
46
|
+
from binboost import BinBoost
|
|
47
|
+
|
|
48
|
+
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
49
|
+
y = np.array([1, 0, 1, 0])
|
|
50
|
+
|
|
51
|
+
model = BinBoost(n_estimators=100, learning_rate=0.2, max_rule_length=4)
|
|
52
|
+
model.fit(X, y)
|
|
53
|
+
|
|
54
|
+
print(model.predict(X))
|
|
55
|
+
print(model.predict_proba(X))
|
|
56
|
+
print(model.rules_)
|
|
57
|
+
|
|
58
|
+
import pandas as pd
|
|
59
|
+
print(pd.DataFrame(model.rule_summary_))
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
## Penggunaan Multi-Kelas (>2 kelas)
|
|
63
|
+
Untuk label dengan lebih dari 2 kelas yang saling eksklusif, gunakan `OneVsRestBinBoost` — hyperparameter yang diteruskan persis sama dengan `BinBoost`:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
from binboost.multiclass import OneVsRestBinBoost
|
|
67
|
+
|
|
68
|
+
model = OneVsRestBinBoost(
|
|
69
|
+
mode='multiclass',
|
|
70
|
+
n_estimators=100,
|
|
71
|
+
learning_rate=0.2,
|
|
72
|
+
max_rule_length=4,
|
|
73
|
+
lambda0=3.0,
|
|
74
|
+
)
|
|
75
|
+
model.fit(X, y) # y boleh berisi >2 kelas, mis. array string atau integer
|
|
76
|
+
|
|
77
|
+
print(model.predict(X))
|
|
78
|
+
print(model.predict_proba(X)) # dinormalisasi supaya tiap baris berjumlah 1
|
|
79
|
+
print(pd.DataFrame(model.rule_summary_)) # aturan per kelas, kolom 'kelas' menandai sub-model asalnya
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
`mode='multilabel'` tersedia untuk kasus di mana satu sampel bisa punya lebih dari satu label positif sekaligus dimana probabilitas tidak dinormalisasi, tiap kelas diputuskan independen.
|
|
83
|
+
|
|
84
|
+
## Hyperparameter Utama
|
|
85
|
+
|
|
86
|
+
| Parameter | Bawaan | Keterangan |
|
|
87
|
+
|---|---|---|
|
|
88
|
+
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
89
|
+
| `learning_rate` | 0.2 | Faktor penyusutan tiap aturan |
|
|
90
|
+
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
91
|
+
| `max_rule_length` | 4 | Jumlah maksimum fitur dalam satu aturan |
|
|
92
|
+
| `operators` | `['AND','OR']` | Operator logika yang digunakan ubah `use_xor=True` untuk mengaktifkan XOR |
|
|
93
|
+
| `beam_width` | 5 | Lebar beam search |
|
|
94
|
+
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
95
|
+
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
96
|
+
| `feature_selection_threshold` | 0.0 | Ambang batas persentil gain fitur untuk seleksi kandidat (0.0 = semua fitur) |
|
|
97
|
+
| `min_samples_rule` | 5 | Jumlah/fraksi minimum sampel yang harus memenuhi sebuah aturan |
|
|
98
|
+
| `lambda0` | 3.0 | Konstanta regularisasi Newton adaptif |
|
|
99
|
+
| `gamma` | 0.0 | Ambang minimum gain untuk menerima sebuah aturan (0.0 = tanpa pruning, nonaktif secara bawaan) |
|
|
100
|
+
|
|
101
|
+
Fitur negasi (NOT) aktif otomatis pada pencarian aturan dan bukan hyperparameter yang bisa diatur lewat `BinBoost()` — ini bagian tetap dari mekanisme beam search.
|
|
102
|
+
|
|
103
|
+
## Fitur yang Didukung
|
|
104
|
+
- Fitur biner (0/1): langsung diproses
|
|
105
|
+
- Fitur numerik (int/float): dibinarisasi otomatis per iterasi
|
|
106
|
+
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
107
|
+
|
|
108
|
+
```python
|
|
109
|
+
from sklearn.preprocessing import OneHotEncoder
|
|
110
|
+
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
111
|
+
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
115
|
+
|
|
116
|
+
## Atribut Setelah Pelatihan
|
|
117
|
+
```python
|
|
118
|
+
model.rules_ # daftar teks aturan termasuk NOT jika ada
|
|
119
|
+
model.rule_weights_ # bobot dua sisi [w0, w1] tiap aturan (hasil Newton + regularisasi)
|
|
120
|
+
model.feature_importances_ # skor kepentingan fitur
|
|
121
|
+
model.train_score_ # loss per iterasi
|
|
122
|
+
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
123
|
+
model.n_rules_ # jumlah aturan aktif
|
|
124
|
+
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
125
|
+
model.thresholds_ # riwayat threshold biner per fitur numerik, per iterasi
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
## Lisensi
|
|
129
|
+
MIT
|
|
@@ -2,7 +2,7 @@ from setuptools import setup, find_packages
|
|
|
2
2
|
|
|
3
3
|
setup(
|
|
4
4
|
name="binboost",
|
|
5
|
-
version="0.2.
|
|
5
|
+
version="0.2.3",
|
|
6
6
|
author="Rangga Wahyu Pratama",
|
|
7
7
|
description="Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner",
|
|
8
8
|
long_description=open("README.md", encoding="utf-8").read(),
|
binboost-0.2.1/PKG-INFO
DELETED
|
@@ -1,101 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: binboost
|
|
3
|
-
Version: 0.2.1
|
|
4
|
-
Summary: Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner
|
|
5
|
-
Author: Rangga Wahyu Pratama
|
|
6
|
-
License: MIT
|
|
7
|
-
Keywords: gradient boosting,logical rules,binary features,interpretable machine learning
|
|
8
|
-
Classifier: Programming Language :: Python :: 3
|
|
9
|
-
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
-
Classifier: Operating System :: OS Independent
|
|
11
|
-
Classifier: Intended Audience :: Science/Research
|
|
12
|
-
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
13
|
-
Requires-Python: >=3.8
|
|
14
|
-
Description-Content-Type: text/markdown
|
|
15
|
-
License-File: LICENSE
|
|
16
|
-
Requires-Dist: numpy>=1.21.0
|
|
17
|
-
Requires-Dist: pandas>=1.3.0
|
|
18
|
-
Dynamic: author
|
|
19
|
-
Dynamic: license-file
|
|
20
|
-
Dynamic: requires-python
|
|
21
|
-
|
|
22
|
-
# BinBoost
|
|
23
|
-
|
|
24
|
-
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner**
|
|
25
|
-
|
|
26
|
-
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
27
|
-
|
|
28
|
-
## Kebaruan Utama
|
|
29
|
-
|
|
30
|
-
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
31
|
-
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti `(A AND B)` atau `(C OR D)`
|
|
32
|
-
|
|
33
|
-
## Instalasi
|
|
34
|
-
|
|
35
|
-
```bash
|
|
36
|
-
pip install binboost
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
## Penggunaan Dasar
|
|
40
|
-
|
|
41
|
-
```python
|
|
42
|
-
import numpy as np
|
|
43
|
-
from binboost import BinBoost
|
|
44
|
-
|
|
45
|
-
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
46
|
-
y = np.array([1, 0, 1, 0])
|
|
47
|
-
|
|
48
|
-
model = BinBoost(n_estimators=50, learning_rate=0.1, max_rule_length=2)
|
|
49
|
-
model.fit(X, y)
|
|
50
|
-
|
|
51
|
-
print(model.predict(X))
|
|
52
|
-
print(model.predict_proba(X))
|
|
53
|
-
print(model.rules_)
|
|
54
|
-
|
|
55
|
-
import pandas as pd
|
|
56
|
-
print(pd.DataFrame(model.rule_summary_))
|
|
57
|
-
```
|
|
58
|
-
|
|
59
|
-
## Hyperparameter Utama
|
|
60
|
-
|
|
61
|
-
| Parameter | Bawaan | Keterangan |
|
|
62
|
-
|---|---|---|
|
|
63
|
-
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
64
|
-
| `learning_rate` | 0.1 | Faktor penyusutan tiap aturan |
|
|
65
|
-
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
66
|
-
| `max_rule_length` | 2 | Jumlah maksimum fitur dalam satu aturan |
|
|
67
|
-
| `operators` | `['AND','OR']` | Operator logika yang digunakan |
|
|
68
|
-
| `beam_width` | 5 | Lebar beam search |
|
|
69
|
-
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
70
|
-
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
71
|
-
| `feature_selection_threshold` | 0.01 | Ambang batas seleksi fitur berbasis gradien |
|
|
72
|
-
|
|
73
|
-
## Fitur yang Didukung
|
|
74
|
-
|
|
75
|
-
- Fitur biner (0/1): langsung diproses
|
|
76
|
-
- Fitur numerik (int/float): dibinarisasi otomatis
|
|
77
|
-
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
78
|
-
|
|
79
|
-
```python
|
|
80
|
-
from sklearn.preprocessing import OneHotEncoder
|
|
81
|
-
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
82
|
-
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
86
|
-
|
|
87
|
-
## Atribut Setelah Pelatihan
|
|
88
|
-
|
|
89
|
-
```python
|
|
90
|
-
model.rules_ # daftar teks aturan
|
|
91
|
-
model.rule_weights_ # bobot setiap aturan
|
|
92
|
-
model.feature_importances_ # skor kepentingan fitur
|
|
93
|
-
model.train_score_ # loss per iterasi
|
|
94
|
-
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
95
|
-
model.n_rules_ # jumlah aturan aktif
|
|
96
|
-
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
## Lisensi
|
|
100
|
-
|
|
101
|
-
MIT
|
binboost-0.2.1/README.md
DELETED
|
@@ -1,80 +0,0 @@
|
|
|
1
|
-
# BinBoost
|
|
2
|
-
|
|
3
|
-
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner**
|
|
4
|
-
|
|
5
|
-
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
6
|
-
|
|
7
|
-
## Kebaruan Utama
|
|
8
|
-
|
|
9
|
-
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
10
|
-
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti `(A AND B)` atau `(C OR D)`
|
|
11
|
-
|
|
12
|
-
## Instalasi
|
|
13
|
-
|
|
14
|
-
```bash
|
|
15
|
-
pip install binboost
|
|
16
|
-
```
|
|
17
|
-
|
|
18
|
-
## Penggunaan Dasar
|
|
19
|
-
|
|
20
|
-
```python
|
|
21
|
-
import numpy as np
|
|
22
|
-
from binboost import BinBoost
|
|
23
|
-
|
|
24
|
-
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
25
|
-
y = np.array([1, 0, 1, 0])
|
|
26
|
-
|
|
27
|
-
model = BinBoost(n_estimators=50, learning_rate=0.1, max_rule_length=2)
|
|
28
|
-
model.fit(X, y)
|
|
29
|
-
|
|
30
|
-
print(model.predict(X))
|
|
31
|
-
print(model.predict_proba(X))
|
|
32
|
-
print(model.rules_)
|
|
33
|
-
|
|
34
|
-
import pandas as pd
|
|
35
|
-
print(pd.DataFrame(model.rule_summary_))
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
## Hyperparameter Utama
|
|
39
|
-
|
|
40
|
-
| Parameter | Bawaan | Keterangan |
|
|
41
|
-
|---|---|---|
|
|
42
|
-
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
43
|
-
| `learning_rate` | 0.1 | Faktor penyusutan tiap aturan |
|
|
44
|
-
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
45
|
-
| `max_rule_length` | 2 | Jumlah maksimum fitur dalam satu aturan |
|
|
46
|
-
| `operators` | `['AND','OR']` | Operator logika yang digunakan |
|
|
47
|
-
| `beam_width` | 5 | Lebar beam search |
|
|
48
|
-
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
49
|
-
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
50
|
-
| `feature_selection_threshold` | 0.01 | Ambang batas seleksi fitur berbasis gradien |
|
|
51
|
-
|
|
52
|
-
## Fitur yang Didukung
|
|
53
|
-
|
|
54
|
-
- Fitur biner (0/1): langsung diproses
|
|
55
|
-
- Fitur numerik (int/float): dibinarisasi otomatis
|
|
56
|
-
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
57
|
-
|
|
58
|
-
```python
|
|
59
|
-
from sklearn.preprocessing import OneHotEncoder
|
|
60
|
-
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
61
|
-
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
62
|
-
```
|
|
63
|
-
|
|
64
|
-
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
65
|
-
|
|
66
|
-
## Atribut Setelah Pelatihan
|
|
67
|
-
|
|
68
|
-
```python
|
|
69
|
-
model.rules_ # daftar teks aturan
|
|
70
|
-
model.rule_weights_ # bobot setiap aturan
|
|
71
|
-
model.feature_importances_ # skor kepentingan fitur
|
|
72
|
-
model.train_score_ # loss per iterasi
|
|
73
|
-
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
74
|
-
model.n_rules_ # jumlah aturan aktif
|
|
75
|
-
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
## Lisensi
|
|
79
|
-
|
|
80
|
-
MIT
|
|
@@ -1,101 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.4
|
|
2
|
-
Name: binboost
|
|
3
|
-
Version: 0.2.1
|
|
4
|
-
Summary: Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner
|
|
5
|
-
Author: Rangga Wahyu Pratama
|
|
6
|
-
License: MIT
|
|
7
|
-
Keywords: gradient boosting,logical rules,binary features,interpretable machine learning
|
|
8
|
-
Classifier: Programming Language :: Python :: 3
|
|
9
|
-
Classifier: License :: OSI Approved :: MIT License
|
|
10
|
-
Classifier: Operating System :: OS Independent
|
|
11
|
-
Classifier: Intended Audience :: Science/Research
|
|
12
|
-
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
13
|
-
Requires-Python: >=3.8
|
|
14
|
-
Description-Content-Type: text/markdown
|
|
15
|
-
License-File: LICENSE
|
|
16
|
-
Requires-Dist: numpy>=1.21.0
|
|
17
|
-
Requires-Dist: pandas>=1.3.0
|
|
18
|
-
Dynamic: author
|
|
19
|
-
Dynamic: license-file
|
|
20
|
-
Dynamic: requires-python
|
|
21
|
-
|
|
22
|
-
# BinBoost
|
|
23
|
-
|
|
24
|
-
**Gradient Boosting Berbasis Aturan Logika Adaptif untuk Fitur Biner**
|
|
25
|
-
|
|
26
|
-
BinBoost adalah algoritma klasifikasi gradient boosting yang membangun ensemble aturan logika murni (AND, OR, XOR) dengan binarisasi fitur numerik adaptif berbasis gradien pada setiap iterasi boosting. Setiap weak learner berupa aturan yang dapat dibaca langsung oleh manusia tanpa memerlukan alat bantu penjelasan pasca-pelatihan.
|
|
27
|
-
|
|
28
|
-
## Kebaruan Utama
|
|
29
|
-
|
|
30
|
-
- **Binarisasi adaptif berbasis gradien**: nilai ambang batas fitur numerik dicari per iterasi untuk memaksimalkan korelasi dengan gradien saat ini
|
|
31
|
-
- **Ensemble aturan logika murni**: tidak ada pohon keputusan, setiap weak learner adalah aturan seperti `(A AND B)` atau `(C OR D)`
|
|
32
|
-
|
|
33
|
-
## Instalasi
|
|
34
|
-
|
|
35
|
-
```bash
|
|
36
|
-
pip install binboost
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
## Penggunaan Dasar
|
|
40
|
-
|
|
41
|
-
```python
|
|
42
|
-
import numpy as np
|
|
43
|
-
from binboost import BinBoost
|
|
44
|
-
|
|
45
|
-
X = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0], [0, 0, 1]], dtype=float)
|
|
46
|
-
y = np.array([1, 0, 1, 0])
|
|
47
|
-
|
|
48
|
-
model = BinBoost(n_estimators=50, learning_rate=0.1, max_rule_length=2)
|
|
49
|
-
model.fit(X, y)
|
|
50
|
-
|
|
51
|
-
print(model.predict(X))
|
|
52
|
-
print(model.predict_proba(X))
|
|
53
|
-
print(model.rules_)
|
|
54
|
-
|
|
55
|
-
import pandas as pd
|
|
56
|
-
print(pd.DataFrame(model.rule_summary_))
|
|
57
|
-
```
|
|
58
|
-
|
|
59
|
-
## Hyperparameter Utama
|
|
60
|
-
|
|
61
|
-
| Parameter | Bawaan | Keterangan |
|
|
62
|
-
|---|---|---|
|
|
63
|
-
| `n_estimators` | 100 | Jumlah iterasi boosting |
|
|
64
|
-
| `learning_rate` | 0.1 | Faktor penyusutan tiap aturan |
|
|
65
|
-
| `loss` | `'logistic'` | Fungsi loss: `'logistic'`, `'focal'`, `'poly'` |
|
|
66
|
-
| `max_rule_length` | 2 | Jumlah maksimum fitur dalam satu aturan |
|
|
67
|
-
| `operators` | `['AND','OR']` | Operator logika yang digunakan |
|
|
68
|
-
| `beam_width` | 5 | Lebar beam search |
|
|
69
|
-
| `binarize_strategy` | `'gradient'` | Strategi binarisasi: `'gradient'`, `'quantile'`, `'uniform'`, `'kmeans'` |
|
|
70
|
-
| `subsample` | 0.8 | Fraksi data per iterasi |
|
|
71
|
-
| `feature_selection_threshold` | 0.01 | Ambang batas seleksi fitur berbasis gradien |
|
|
72
|
-
|
|
73
|
-
## Fitur yang Didukung
|
|
74
|
-
|
|
75
|
-
- Fitur biner (0/1): langsung diproses
|
|
76
|
-
- Fitur numerik (int/float): dibinarisasi otomatis
|
|
77
|
-
- Fitur kategorikal 3+ kelas: wajib OneHotEncode terlebih dahulu
|
|
78
|
-
|
|
79
|
-
```python
|
|
80
|
-
from sklearn.preprocessing import OneHotEncoder
|
|
81
|
-
enc = OneHotEncoder(sparse_output=False, drop='first')
|
|
82
|
-
X_encoded = enc.fit_transform(X[['kolom_kategorikal']])
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
BinBoost secara otomatis mendeteksi kelompok fitur OneHotEncoding dan mencegah aturan yang tidak masuk akal seperti `(Warna_Merah AND Warna_Biru)`.
|
|
86
|
-
|
|
87
|
-
## Atribut Setelah Pelatihan
|
|
88
|
-
|
|
89
|
-
```python
|
|
90
|
-
model.rules_ # daftar teks aturan
|
|
91
|
-
model.rule_weights_ # bobot setiap aturan
|
|
92
|
-
model.feature_importances_ # skor kepentingan fitur
|
|
93
|
-
model.train_score_ # loss per iterasi
|
|
94
|
-
model.rule_summary_ # ringkasan lengkap atau konversi ke DataFrame
|
|
95
|
-
model.n_rules_ # jumlah aturan aktif
|
|
96
|
-
model.feature_usage_ # frekuensi penggunaan tiap fitur
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
## Lisensi
|
|
100
|
-
|
|
101
|
-
MIT
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|