gaca 1.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
gaca-1.2.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2025 henrivuorinen
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
gaca-1.2.0/PKG-INFO ADDED
@@ -0,0 +1,541 @@
1
+ Metadata-Version: 2.4
2
+ Name: gaca
3
+ Version: 1.2.0
4
+ Summary: Gravitational Accretion Clustering: physics-inspired clustering and anomaly detection for large tables
5
+ Author: Henri J. Vuorinen
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/henrivuorinen/GACA
8
+ Project-URL: Repository, https://github.com/henrivuorinen/GACA
9
+ Project-URL: Issues, https://github.com/henrivuorinen/GACA/issues
10
+ Project-URL: Changelog, https://github.com/henrivuorinen/GACA/blob/main/CHANGELOG.md
11
+ Keywords: clustering,anomaly detection,outlier detection,mean shift,density clustering,unsupervised learning,streaming
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Science/Research
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Topic :: Scientific/Engineering :: Information Analysis
18
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
19
+ Classifier: Topic :: Scientific/Engineering :: Astronomy
20
+ Requires-Python: >=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: numpy>=1.24
24
+ Requires-Dist: scipy>=1.10
25
+ Requires-Dist: scikit-learn>=1.3
26
+ Requires-Dist: pandas>=2.0
27
+ Provides-Extra: cli
28
+ Provides-Extra: experiments
29
+ Requires-Dist: matplotlib>=3.7; extra == "experiments"
30
+ Requires-Dist: hdbscan>=0.8; extra == "experiments"
31
+ Requires-Dist: lightgbm>=4.0; extra == "experiments"
32
+ Requires-Dist: psutil>=5.9; extra == "experiments"
33
+ Requires-Dist: pyarrow>=14; extra == "experiments"
34
+ Provides-Extra: dev
35
+ Requires-Dist: pytest>=7; extra == "dev"
36
+ Requires-Dist: build; extra == "dev"
37
+ Requires-Dist: twine; extra == "dev"
38
+ Dynamic: license-file
39
+
40
+ # GACA: Gravitational Accretion Clustering Algorithm
41
+
42
+ GACA is a density-based clustering method that treats every data point as a
43
+ particle with mass. Nearby mass attracts, particles drift towards dense regions,
44
+ and particles that meet merge and keep their combined mass. What is left are the
45
+ **Suns**: the cluster centres. You never specify the number of clusters. It
46
+ emerges from the dynamics and is controlled by one bandwidth parameter, γ.
47
+
48
+ GACA is designed for large tables. The expensive simulation runs once on a small
49
+ random sample (the coreset), and the full dataset is then streamed through the
50
+ resulting Suns in batches. In the thesis experiments this clustered **10 million
51
+ records in about 23 seconds on a laptop, with constant memory**.
52
+
53
+ This repository contains the reference implementation from the master's thesis
54
+ *The Gravitational Accretion Clustering Algorithm (GACA): Scalable,
55
+ Physics-Inspired Clustering for Large Datasets*, plus several improvements made
56
+ after it (see [Improvements since the thesis](#improvements-since-the-thesis)).
57
+ The thesis behaviour is still available with
58
+ `GACA(assignment='newton', method='barnes_hut')`, and every script in
59
+ `experiments/` uses it.
60
+
61
+ ![Solar Genesis: 3,000 particles condensing into 7 Suns](https://raw.githubusercontent.com/henrivuorinen/GACA/main/figures/accretion_evolution.png)
62
+
63
+ ## How it works
64
+
65
+ **Phase 1: Solar Genesis (structure discovery on a coreset).** A uniform sample of
66
+ `sample_size` rows (5,000 by default) starts as particles of mass 1. Each
67
+ iteration:
68
+
69
+ 1. **Transport.** Every particle moves a damped step (η) towards the mass-weighted
70
+ Gaussian average of its surroundings. This is exactly a blurring mean-shift
71
+ step: the displacement equals `(1 / 2γ) ∇ log KDE(x)`.
72
+ 2. **Condensation.** Particles closer than ε are linked, and each connected
73
+ component is replaced by one particle at its centre of mass, carrying the
74
+ summed mass. Mass is conserved exactly.
75
+
76
+ The Gaussian sum is evaluated exactly in vectorised blocks (`method='exact'`,
77
+ the default). The thesis used a Barnes-Hut tree (`method='barnes_hut'`), which is
78
+ asymptotically cheaper but in interpreted Python is 90 to 140 times slower on
79
+ the coresets GACA uses, and only approximate. Genesis stops at a stable plateau
80
+ or at the iteration cap. The surviving particles are the Suns, and every
81
+ coreset row remembers which Sun it condensed into.
82
+
83
+ **Phase 2: Particle Accretion (assignment of every row).** Each row is assigned
84
+ to the Sun whose coreset members exert the strongest Gaussian pull on it,
85
+ `Σ_{j ∈ Sun k} exp(−γ‖x − z_j‖²)`, summed over its 32 nearest coreset members.
86
+ This follows the shape the dynamics found rather than cutting space into
87
+ Voronoi cells. It is O(N log n_c) time and works batch by batch, so memory
88
+ depends on the batch size rather than on N. (`assignment='newton'` restores
89
+ the thesis rule, `M_k / d²`.)
90
+
91
+ **Lone Suns (anomalies).** A point far from everything feels essentially no
92
+ Gaussian pull, never moves and never merges. It survives as a Sun of mass 1: a
93
+ Lone Sun. Assignment applies the same test to every streamed row: if the total
94
+ pull on it is below `kappa_`, the row would not have moved had it been in the
95
+ coreset (thesis Prop. 3.8), so it becomes a new Lone Sun, or joins an earlier
96
+ Lone Sun that pulls it. Anomalies therefore no longer need to be in the coreset
97
+ to be isolated.
98
+
99
+ **Saddle linking (optional, `link_tau`).** At a fine bandwidth a curved or
100
+ elongated cluster splits into a chain of Suns. With `link_tau` set, two Suns
101
+ are joined when the density along the best edge between their members stays
102
+ above `link_tau` times the lower of their two peaks, so the decision rests on
103
+ the mass lying between the Suns rather than the distance between them. Lone
104
+ Suns are never joined.
105
+
106
+ ## Installation
107
+
108
+ ```bash
109
+ pip install gaca
110
+ ```
111
+
112
+ That installs the package and the `gaca` command (numpy, scipy, scikit-learn
113
+ and pandas are pulled in). Python 3.10 or newer.
114
+
115
+ For the latest development version, or to run the tests and experiments:
116
+
117
+ ```bash
118
+ pip install "git+https://github.com/henrivuorinen/GACA.git" # latest from GitHub
119
+
120
+ git clone https://github.com/henrivuorinen/GACA.git # or work on the code
121
+ cd GACA
122
+ pip install -e ".[dev]" # + pytest, build, twine
123
+ pip install -e ".[experiments]" # + matplotlib, hdbscan, lightgbm, ... for experiments/
124
+ ```
125
+
126
+ ## Any table: AutoGACA and the `gaca` command
127
+
128
+ The algorithm only needs points in a space, but someone has to decide how to
129
+ turn a table into that space and which bandwidth to use. `AutoGACA` makes those
130
+ decisions from the data and records each one, so a dataset can be clustered
131
+ without knowing anything about it in advance:
132
+
133
+ ```bash
134
+ gaca run galaxies.csv # writes galaxies_gaca/
135
+ gaca run galaxies.csv --exclude objid,ra,dec # leave columns out
136
+ gaca run huge.csv --chunksize 200000 # files larger than memory
137
+ ```
138
+
139
+ The output folder holds `labels.csv` (the input rows plus `gaca_cluster`,
140
+ `gaca_anomaly`, `gaca_anomaly_group` and `gaca_anomaly_score`), a self-contained
141
+ `report.html` and `summary.json`. The report describes each cluster in terms of
142
+ the original columns, lists the anomalies and the columns that make them
143
+ unusual, shows how the resolution was chosen and which other resolutions were
144
+ also stable, and lists every preprocessing decision.
145
+
146
+ From Python:
147
+
148
+ ```python
149
+ from gaca import AutoGACA
150
+
151
+ auto = AutoGACA().fit(df) # DataFrame, array, or path to a CSV
152
+ print(auto.summary())
153
+ auto.result_ # per-row cluster, anomaly flag, group, score
154
+ auto.report("report.html", data=df)
155
+ new = auto.assign(df_new) # same preprocessing, same clusters
156
+ ```
157
+
158
+ What it decides, and how:
159
+
160
+ | Step | Rule |
161
+ |---|---|
162
+ | Columns | Numeric columns are used. Text, identifiers (unique integers named like `id` or strictly increasing), constant columns and columns more than 50% missing are left out. `columns=` / `exclude=` override. |
163
+ | Missing values | Median of the column. |
164
+ | Transforms | `log1p` for count- or flux-like columns: non-negative, skewed, and spanning more than an order of magnitude. Other skewed columns are left alone, since a log would also pull genuine outliers back towards the data. |
165
+ | Scaling | Median and interquartile range, so outliers do not compress everything else. `scale='none'` (`--scale none`) keeps the values as they are, for data whose distances already mean something, such as positions in physical units. |
166
+ | Dimensions | Unchanged up to `max_dims` columns (default 10). Above that, PCA to between 5 and `max_dims` components, the number set by the noise floor of the singular values (Gavish and Donoho). |
167
+ | γ | Swept as c / (median squared distance between rows). At each value GACA is fitted on four subsamples. γ is the middle of the most stable stretch of the longest plateau in the cluster count. A split only counts if the subsamples agree (adjusted Rand index ≥ 0.8). Otherwise the answer is one group. |
168
+ | Anomalies | Rows that no cluster pulls, plus clusters smaller than `min_cluster_size` (10 rows). Score = −log10 of the pull relative to a typical row; above 3 is anomalous. |
169
+ | Hierarchy | Every stable resolution of the sweep is kept, fitted on the same coreset and nested into a tree: each fine cluster sits under the coarse cluster holding most of its particles, and each row is assigned within its own branch. `gaca_cluster` is the level the rule above picks; `gaca_level_1 … gaca_level_K` hold all levels as path-style names (`0`, `0.2`, `0.2.1`). `hierarchy=False` turns it off. |
170
+
171
+ How it did on data it was not tuned on (ARI against known labels; *best γ* is
172
+ the best that any γ in the sweep achieved):
173
+
174
+ | Data | Rows × columns | AutoGACA | Best γ | Notes |
175
+ |---|---|---|---|---|
176
+ | Wine | 178 × 13 | 0.79 | 0.82 | 3 clusters, PCA to 5-D |
177
+ | Iris | 150 × 4 | 0.54 | 0.55 | two of the species overlap; density methods see 2 groups |
178
+ | Breast cancer | 569 × 30 | 0.36 | 0.50 | |
179
+ | Blobs + planted anomalies | 20,042 × 5 | 0.99 | 0.99 | 42/42 anomalies, 0.1% false positives |
180
+ | 6 clusters in 30-D, 25 columns of noise | 6,000 × 30 | 0.90 | 0.89 | PCA to 5-D |
181
+ | Two moons, `link_tau=0.6` | 4,000 × 2 | 1.00 | | |
182
+ | Concentric circles, `link_tau=0.6` | 4,000 × 2 | 0.93 | | |
183
+ | Handwritten digits | 1,797 × 64 | 0.17 | 0.24 | not density-separated in a few dimensions |
184
+ | One Gaussian (no clusters) | 5,000 × 2 to 10 | 1 cluster | | correctly reports no structure and no anomalies |
185
+
186
+ `link_tau` is not on by default: it joins curved clusters but also merges
187
+ clusters that overlap (the 30-D case drops to 0.15 with it).
188
+
189
+ ### The cluster hierarchy
190
+
191
+ Real data often has structure at several scales, and one automatic resolution
192
+ then has to pick between a coarse and a fine view that are almost equally
193
+ stable. AutoGACA keeps all of them. On the SDSS objects (no labels used), the
194
+ tree runs from 3 clusters (the main population, quasars, cool M dwarfs) to 10
195
+ (distant luminous galaxies, white dwarfs, low-redshift quasars and warmer M
196
+ dwarfs separate out) to 17, where the quasars form a ladder of 99 to 100% pure
197
+ clusters by redshift (z ≈ 0.9, 1.2, 1.6 and 2.0). The report shows the tree
198
+ with what sets each branch apart from its parent.
199
+
200
+ Two details make the tree reliable:
201
+
202
+ - The bandwidth sweep reaches down to the local scale (a kernel about three
203
+ nearest-neighbour distances wide), not just to a fixed fraction of the
204
+ typical distance. Otherwise groups made of subgroups never get resolved.
205
+ - Between equally long, equally stable plateaus the finer one is used for
206
+ `gaca_cluster`, since the coarser ones remain in the hierarchy.
207
+
208
+ One caution: robust scaling divides each column by its own spread. A column
209
+ whose values form far-apart groups therefore gets compressed, which can merge
210
+ the subgroups along it. When all columns share a unit, `scale='none'` avoids
211
+ this.
212
+
213
+ To try it on real astronomy data, `examples/sdss/` downloads two public Sloan
214
+ Digital Sky Survey tables (object properties, and galaxy positions around the
215
+ Coma cluster) and walks through both runs; see
216
+ [examples/sdss/README.md](https://github.com/henrivuorinen/GACA/blob/main/examples/sdss/README.md).
217
+
218
+ ### Why `max_dims` can be 10 or more
219
+
220
+ The thesis put the ceiling at about 15 dimensions, for two reasons: the
221
+ Barnes-Hut tree degrades, and the Gaussian kernel loses contrast. The first is
222
+ gone with the exact backend. The second was made worse in the thesis experiment
223
+ by keeping γ = 1 at every dimension: in 45-D, squared distances are around 90,
224
+ every pull is about e⁻⁹⁰, and every point becomes a singleton. With γ scaled to
225
+ the data, six clusters were recovered at ARI ≥ 0.98 from 10 to 50 dimensions
226
+ when every column carries signal. When most columns are noise the limit is
227
+ real: raw data held up to about 15 dimensions (ARI 0.86), fell to 0.17 at 30,
228
+ and PCA restored it (0.89 at 5 components). Hence the default: use the columns
229
+ as they are up to 10, and project above that. Raise `max_dims` when you know the
230
+ columns are informative.
231
+
232
+ ## Benchmark against established methods
233
+
234
+ `experiments/benchmark_suite.py` compares AutoGACA with HDBSCAN, Isolation
235
+ Forest, LOF and K-Means. Every method gets the same preprocessed input, and
236
+ nothing is tuned per dataset with the labels.
237
+
238
+ **Clustering** (adjusted Rand index; HDBSCAN with minimum cluster size 1% of
239
+ rows, K-Means given the true k as a reference):
240
+
241
+ | Data | GACA | GACA + link | HDBSCAN | K-Means (true k) |
242
+ |---|---|---|---|---|
243
+ | wine | **0.81** | 0.00 | 0.53 | 0.86 |
244
+ | breast cancer | **0.36** | 0.01 | 0.08 | 0.64 |
245
+ | digits | 0.03 | **0.39** | 0.01 | 0.27 |
246
+ | uneven blobs 5-D | **0.99** | 0.97 | 0.56 | 0.92 |
247
+ | 6 clusters in 30-D noise | **0.90** | 0.15 | 0.13 | 0.90 |
248
+ | varied density | 0.93 | **0.95** | 0.79 | 0.91 |
249
+ | anisotropic | 0.67 | **0.99** | 0.95 | 0.66 |
250
+ | two moons | 0.53 | **1.00** | **1.00** | 0.33 |
251
+ | circles | 0.53 | 0.96 | **1.00** | 0.00 |
252
+ | SDSS objects (class) | 0.48 | **0.61** | **0.61** | 0.36 |
253
+
254
+ GACA is the stronger method on overlapping, Gaussian-like groups, varied
255
+ densities and noisy dimensions. HDBSCAN is stronger on curved shapes, where
256
+ GACA needs `link_tau`; with it, GACA matches or nearly matches HDBSCAN there
257
+ and on the SDSS objects, and is the best method on digits. The two GACA modes
258
+ are complementary, and linking is what fails on overlapping groups (wine drops
259
+ to 0.00 with it). The ARI here is for the single level in `gaca_cluster`; the
260
+ hierarchy often holds a better-matching level as well. GACA also assigns 90
261
+ to 100% of rows, while HDBSCAN leaves up to 68% as noise. Neither a fixed linking ratio nor a significance
262
+ test chose correctly between linking and not linking: GACA's dynamics
263
+ separate overlapping groups that a density-dip test calls a single mode.
264
+
265
+ **Anomaly detection** (average precision; ROC AUC in brackets):
266
+
267
+ | Data | GACA | Isolation Forest | LOF | HDBSCAN |
268
+ |---|---|---|---|---|
269
+ | blobs + planted anomalies | **1.00** (1.00) | 0.97 (1.00) | **1.00** (1.00) | 0.44 (0.81) |
270
+ | company-like + planted | **0.74** (1.00) | 0.62 (1.00) | 0.44 (0.99) | 0.45 (0.76) |
271
+ | breast cancer, 5% malignant | **0.47** (0.95) | 0.34 (0.91) | 0.46 (0.94) | 0.05 (0.53) |
272
+ | SDSS, quasars thinned to 1% | 0.14 (0.96) | **0.28** (0.95) | 0.02 (0.56) | 0.01 (0.52) |
273
+ | SDSS, white dwarfs (~1%) | 0.04 (0.90) | **0.06** (0.94) | 0.01 (0.46) | 0.01 (0.52) |
274
+ | KDD Cup 99 attacks | 0.05 (0.61) | **0.33** (0.80) | 0.03 (0.34) | 0.03 (0.34) |
275
+
276
+ GACA ranks isolated anomalies best and flags precisely: on the planted sets it
277
+ flags 0.3% of rows and catches them all, where Isolation Forest flags 9 to 16%.
278
+ It is weaker when the "anomalies" are a dense group of their own (a thousand
279
+ similar quasars, or the floods of identical connections in KDD Cup, where all
280
+ density methods fail).
281
+
282
+ **Scale** (5-D data; seconds, and peak memory of the process):
283
+
284
+ | Rows | GACA, fixed γ, streamed | AutoGACA | HDBSCAN | Isolation Forest | LOF | K-Means |
285
+ |---|---|---|---|---|---|---|
286
+ | 10,000 | 1.0 s, 519 MB | 9.9 s, 733 MB | 0.3 s, 207 MB | 0.1 s | 0.1 s | 0.0 s |
287
+ | 100,000 | 1.1 s, 542 MB | 10.7 s, 962 MB | 39.8 s, 3.1 GB | 0.3 s | 1.1 s | 0.0 s |
288
+ | 1,000,000 | 2.2 s, 603 MB | 13.8 s, 2.7 GB | did not finish in 10 min | 3.0 s, 469 MB | 29 s, 1.2 GB | 0.2 s |
289
+
290
+ Streamed GACA keeps memory flat. AutoGACA adds a fixed ~10 s for choosing γ,
291
+ and in memory it holds the whole table (the CLI's `--chunksize` mode does not).
292
+ Isolation Forest and K-Means are faster at every size; the speed advantage is
293
+ over density-based clustering.
294
+
295
+ ## Using the algorithm directly
296
+
297
+ ```python
298
+ import numpy as np
299
+ from sklearn.preprocessing import StandardScaler
300
+ from gaca import GACA
301
+
302
+ X = StandardScaler().fit_transform(X_raw) # standardise first
303
+
304
+ model = GACA(gamma_clustering=1.0, random_state=0).fit(X)
305
+ labels = model.assign(X) # Sun index for every row
306
+ anomalies = model.is_lone(labels) # rows in Lone Suns
307
+
308
+ model.suns_ # (k, d) positions of the Suns found by genesis
309
+ model.sun_masses_ # (k,) mass each Sun accumulated from the coreset
310
+ model.n_suns_ # k plus the Lone Suns registered by assign
311
+ ```
312
+
313
+ A runnable demo on synthetic data (five groups plus planted anomalies, compared
314
+ with K-Means):
315
+
316
+ ```bash
317
+ python examples/quickstart.py
318
+ ```
319
+
320
+ ### Streaming data larger than memory
321
+
322
+ `fit` only needs the coreset, and `assign` works on any batch, so a large file
323
+ can be processed chunk by chunk:
324
+
325
+ ```python
326
+ model = GACA(random_state=0).fit(coreset) # e.g. 5,000 sampled rows
327
+ for chunk in pd.read_csv("big.csv", chunksize=100_000):
328
+ labels = model.assign(scaler.transform(chunk[features].values))
329
+ ```
330
+
331
+ Lone Suns registered in one chunk persist, so later copies of the same anomaly
332
+ receive the same label.
333
+
334
+ ### Mixture of Experts
335
+
336
+ Passing a target to `fit` trains one Ridge regression per Sun (logistic
337
+ regression with `task_type='classification'`). `predict` routes each row to its
338
+ Sun's expert. Suns with fewer than `min_expert_size` training rows predict their
339
+ mean. Empty Suns, and rows that no Sun pulls, fall back to a global model.
340
+
341
+ ```python
342
+ model = GACA(random_state=0).fit(X_train, y_train)
343
+ y_pred = model.predict(X_test)
344
+ ```
345
+
346
+ ## Parameters
347
+
348
+ | Parameter | Default | Meaning |
349
+ |---|---|---|
350
+ | `gamma_clustering` | 1.0 | Gaussian bandwidth γ. The main dial: low γ merges everything into one Sun, high γ fragments into many |
351
+ | `epsilon` | `'auto'` | Condensation radius ε. `'auto'` = 0.05/√γ, capped at half the coreset's median nearest-neighbour distance (= 0.05 on the thesis data); a radius larger than the point spacing chains dense structures together |
352
+ | `eta` | 0.5 | Damping η of the transport step |
353
+ | `sample_size` | 5000 | Coreset size for Solar Genesis |
354
+ | `n_iterations` | 20 | Iteration cap for genesis |
355
+ | `theta` | 0.5 | Barnes-Hut opening angle (only with `method='barnes_hut'`) |
356
+ | `method` | `'exact'` | Genesis kernel sum: `'exact'` (vectorised) or `'barnes_hut'` (thesis) |
357
+ | `assignment` | `'pull'` | `'pull'` (member pull, Lone-Sun registration) or `'newton'` (thesis `M/d²`) |
358
+ | `kappa` | 1e-3 | Lone-Sun threshold, relative to the median pull inside the coreset |
359
+ | `link_tau` | None | Saddle-linking ratio; e.g. 0.6 together with γ ≈ 10 for non-convex clusters |
360
+ | `n_neighbors` | 32 | Coreset members used for the pull of each row |
361
+ | `lone_share` | 1e-3 | A genesis Sun with at most this share of the coreset counts as a Lone Sun |
362
+ | `min_expert_size` | 10 | Minimum rows for a Sun to get its own regression expert |
363
+ | `random_state` | None | Seed for the coreset draw |
364
+ | `verbose` | False | Print progress |
365
+
366
+ The lower-level functions are also exported: `solar_genesis` (the simulation on
367
+ its own; `return_members=True` gives the Sun of every input row, `plateau=w` adds
368
+ the plateau stopping rule), `assign` (Newtonian assignment given Suns and
369
+ masses), `kernel_pull`, `saddle_link`, `GACANode` and
370
+ `merge_connected_components`.
371
+
372
+ ## Practical guidance
373
+
374
+ These apply to `GACA` used directly; `AutoGACA` handles the first three.
375
+
376
+ - **Scale your features.** γ is defined relative to the scale of the space.
377
+ - **Mind the dimension.** See [above](#why-max_dims-can-be-10-or-more): with γ
378
+ scaled to the data, 10 or more informative columns work. Many noise columns
379
+ need PCA first.
380
+ - **Choose γ from the plateau.** Sweep γ and pick a value where the Sun count is
381
+ stable (`select_gamma` does this). γ = 1.0 sat in the stable band on the
382
+ thesis data.
383
+ - **Genesis is cheap now.** With the exact backend, genesis takes about 0.6 s on
384
+ a 5,000-point coreset and 28 s on 40,000 (the thesis reports 22.5 s and 370 s).
385
+ - **Non-convex clusters: link.** For crescents, rings or elongated groups, use
386
+ `link_tau=0.6`. Linking joins clusters connected by a dense bridge, so it also
387
+ joins clusters that overlap; leave it off when groups touch.
388
+ - **Anomaly sensitivity.** `kappa` sets how little pull makes a row a Lone Sun.
389
+ Raising γ also makes more rows lone, since the kernel narrows; at very high γ
390
+ a small coreset cannot cover sparse tails and genuine rows start to be
391
+ flagged.
392
+
393
+ ## Results from the thesis
394
+
395
+ | | Result |
396
+ |---|---|
397
+ | Scalability | 10⁷ records: 23.5 s total, 0.81 µs per record for assignment, peak memory flat at 707 MB |
398
+ | vs HDBSCAN | 1,766× faster at 10⁷ records; HDBSCAN is faster below ~3×10⁵ |
399
+ | Structure | Largest-Sun share 0.9203 to 0.9209 while streaming from 10⁵ to 10⁷ records |
400
+ | Fragmentation | 5 effective clusters vs 91 (HDBSCAN) and 94 (Mean-Shift) on 20k records |
401
+ | Anomalies | 92% of sampled planted anomalies isolated as Lone Suns; 100% of unsampled ones flagged by the OOD rule |
402
+ | Robustness | Under heavy feature noise (σ = 2), a K-Means mixture fell to R² = −0.004 while the GACA mixture kept 0.012 |
403
+
404
+ ![Bandwidth sweep: phase transition and downstream utility](https://raw.githubusercontent.com/henrivuorinen/GACA/main/figures/phase_transition.png)
405
+
406
+ ![Runtime versus dataset size](https://raw.githubusercontent.com/henrivuorinen/GACA/main/figures/scalability_comparison.png)
407
+
408
+ ## Improvements since the thesis
409
+
410
+ Measured with `experiments/improvements_benchmark.py` on synthetic data (20,000
411
+ rows, coreset 2,000, γ = 1, 4 seeds, the nine planted anomaly groups of
412
+ Sec. 6.4). *Detected* is the share of anomaly groups whose every copy ends up
413
+ in a Lone Sun; *FP* is the share of genuine rows flagged; *R²* is the Mixture of
414
+ Experts on a fresh test set with new anomalies.
415
+
416
+ | Data | Method | Detected | FP | ARI | R² |
417
+ |---|---|---|---|---|---|
418
+ | blobs | thesis (`M/d²`) | 0.19 | 0.00% | 0.990 | 0.963 |
419
+ | blobs | thesis + OOD95 rule | 1.00 | 4.87% | 0.990 | 0.936 |
420
+ | blobs | **pull (new default)** | 0.97 | **0.00%** | **0.994** | **0.974** |
421
+ | company-like | thesis (`M/d²`) | 0.22 | 0.01% | 0.080 | 0.397 |
422
+ | company-like | thesis + OOD95 rule | 1.00 | 4.89% | 0.080 | 0.395 |
423
+ | company-like | **pull (new default)** | **1.00** | 0.35% | 0.173 | **0.579** |
424
+ | company-like | pull, γ = 5, `link_tau=0.6` | 1.00 | 5.76% | 0.604 | 0.738 |
425
+
426
+ What changed and why:
427
+
428
+ 1. **Exact vectorised genesis.** Same Suns as the Barnes-Hut version (to within a
429
+ unit of mass), no approximation error, and 90 to 140 times faster in this
430
+ implementation. This also settles the "approximation bounds" limitation: the
431
+ default operator is exact.
432
+ 2. **Anomalies no longer depend on the coreset.** The thesis found that an
433
+ unsampled anomaly is never isolated, and that the OOD rule that catches them
434
+ also flags 5% of genuine rows by construction (about 1,000 false alarms per
435
+ 20,000 rows). The pull test catches them at 0 to 0.35% false positives,
436
+ independently of the coreset size, and groups the copies of one anomaly
437
+ into a single Lone Sun.
438
+ 3. **Assignment follows the dynamics.** Routing by the pull of each Sun's
439
+ members instead of `M_k / d²` improved the downstream Mixture of Experts
440
+ (R² 0.40 to 0.58 on the company-like data), and rows that no Sun pulls are
441
+ sent to the global model instead of a wrong local expert.
442
+ 4. **Saddle linking** recovers non-convex clusters that the flat output could
443
+ not express (two moons and concentric circles at γ = 10: ARI 0.2 to 0.5
444
+ without, 0.99 to 1.0 with `link_tau=0.6`, over 4 seeds), and widens the
445
+ usable γ range (five blobs at ARI ≥ 0.94 for γ = 3 to 10, where unlinked
446
+ γ = 10 shatters them into 300 Suns). It is not uniformly safe: at γ = 1 it
447
+ can join adjacent groups, and at γ = 30 it under-merges.
448
+ 5. **Relative Lone-Sun threshold** (`kappa` as a fraction of the median pull) so
449
+ that one setting works across bandwidths, dimensions and coreset sizes.
450
+ 6. **Lone Sun by share, not by mass 1.** Once two copies of a rare group are
451
+ sampled, the thesis definition (mass 1) no longer counts it as an anomaly.
452
+
453
+ Remaining caveats: the 5σ misses in the table are planted points within about
454
+ one unit of genuine data, where the kernel does not consider them separate.
455
+ The real company data was not available for these tests, so the thesis
456
+ numbers should be re-run before the new defaults are trusted on that data.
457
+
458
+ ## Repository layout
459
+
460
+ ```
461
+ gaca/ the package: Solar Genesis and assignment (genesis.py), the
462
+ GACA estimator (model.py), AutoGACA (auto.py), the HTML report
463
+ (report.py) and the gaca command (cli.py)
464
+ tests/ pytest suite (mass conservation, Lone Suns, estimator, AutoGACA, CLI)
465
+ examples/ runnable demo on synthetic data; sdss/ fetches public SDSS
466
+ data and shows AutoGACA on it
467
+ experiments/ the scripts behind every thesis table and figure, plus
468
+ improvements_benchmark.py (thesis vs current defaults)
469
+ figures/ images used in this README
470
+ data/ empty; put your own data here (git-ignored)
471
+ ```
472
+
473
+ ## Reproducing the thesis experiments
474
+
475
+ The thesis used proprietary company data, which is not included. The scripts in
476
+ `experiments/` run on any CSV with the columns below. Run them from the
477
+ repository root; outputs (CSV, PDF, PNG) are written to the current directory.
478
+
479
+ The company-data loader (`experiments/process_data.py`) builds five features:
480
+
481
+ | Feature | Built from |
482
+ |---|---|
483
+ | `log_emp` | `EMPLOYEES_COUNT`, falling back to the midpoint of the `SIZE` band (`1-10`, `11-50`, ...) |
484
+ | `log_reviews` | `TOTAL_GMAPS_REVIEW_COUNT` |
485
+ | `age` | `FOUNDED` (year) |
486
+ | `LATITUDE`, `LONGITUDE` | as is |
487
+
488
+ The regression target is `GMAPS_REVIEWS_AVERAGE`. The characterisation scripts
489
+ also use `NAME`, `GMAPS_PRIMARY_CATEGORY`, `INDUSTRY` and `COUNTRY` when present.
490
+ The dimensionality experiment accepts any wide table of numeric columns.
491
+
492
+ | Script | Thesis result |
493
+ |---|---|
494
+ | `sweep_gamma.py` → `plot_phase_transition.py` | γ phase transition (Sec. 5.1) |
495
+ | `sweep_epsilon.py` | ε sensitivity |
496
+ | `benchmark_clustering.py` | GACA vs K-Means, Mean-Shift, HDBSCAN (Table 5.4) |
497
+ | `final_benchmark_test.py` | Mixture of Experts under feature corruption (RQ3) |
498
+ | `characterize_suns.py`, `plot_suns.py` | What the Suns contain (Sec. 5.2.1) |
499
+ | `anomaly_recovery.py` | Planted-anomaly recovery through the coreset |
500
+ | `convergence_check.py` | Long-run convergence and plateaus |
501
+ | `dimensionality_experiment.py` → `plot_dimensionality.py` | Curse of dimensionality and PCA rescue (Sec. 4.3.2) |
502
+ | `run_scalability_test.py` | 10⁷-record streaming run (Ch. 6) |
503
+ | `scalability_comparison.py` → `plot_scalability_comparison.py` | Runtime vs baselines (Ch. 6) |
504
+ | `coreset_sweep.py` → `plot_coreset.py` | Coreset-size sensitivity (Ch. 6) |
505
+ | `plot_accretion_evolution.py` | The accretion figures |
506
+ | `make_tables.py` | Writes every thesis table as LaTeX from the result CSVs |
507
+
508
+ Example:
509
+
510
+ ```bash
511
+ python experiments/sweep_gamma.py
512
+ python experiments/plot_phase_transition.py sweep_results_eta0.5.csv
513
+ ```
514
+
515
+ ## Limitations
516
+
517
+ - The Barnes-Hut acceptance rule (`method='barnes_hut'`) has no uniform error
518
+ bound for a Gaussian kernel. The default exact backend avoids it, at O(n²) per
519
+ iteration, which is fine up to coresets of a few tens of thousands.
520
+ - There is no global optimality guarantee, and the result is the state at the
521
+ stopping time (the stopping rule is part of the algorithm).
522
+ - A uniform coreset can miss rare structure. Isolated anomalies are recovered at
523
+ assignment time, but a small genuine cluster missed by the coreset becomes a
524
+ set of Lone Suns rather than a regular Sun.
525
+ - Only low-dimensional inputs are practical; use a projection first.
526
+
527
+ ## Citation
528
+
529
+ ```bibtex
530
+ @mastersthesis{vuorinen2026gaca,
531
+ author = {Vuorinen, Henri J.},
532
+ title = {The Gravitational Accretion Clustering Algorithm (GACA): Scalable,
533
+ Physics-Inspired Clustering for Large Datasets},
534
+ school = {Aalto University},
535
+ year = {2026}
536
+ }
537
+ ```
538
+
539
+ ## License
540
+
541
+ MIT, see [LICENSE](https://github.com/henrivuorinen/GACA/blob/main/LICENSE).