specunet-pkg 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,346 @@
1
+ Metadata-Version: 2.4
2
+ Name: specunet_pkg
3
+ Version: 1.0.0
4
+ Summary: An implementation of a Conventional UNet for spectral image denoising.
5
+ Home-page: https://github.com/switfluors/SpecUNet.git
6
+ Author: Obblivignes KanchanadeviVenkataraman, Hongjing Mao, Dr. Dongkuan Xu, Dr. Yang Zhang, Dr. Caroline Laplante
7
+ Author-email: switfluors@gmail.com
8
+ License: "MIT"
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: Operating System :: OS Independent
11
+ Classifier: Intended Audience :: Science/Research
12
+ Classifier: Topic :: Scientific/Engineering :: Image Processing
13
+ Requires-Python: >=3.8
14
+ Description-Content-Type: text/markdown
15
+ Requires-Dist: numpy
16
+ Requires-Dist: matplotlib
17
+ Requires-Dist: h5py
18
+ Requires-Dist: hdf5storage
19
+ Requires-Dist: pillow
20
+ Requires-Dist: scipy
21
+ Requires-Dist: tqdm
22
+ Requires-Dist: onnx
23
+ Requires-Dist: onnxscript
24
+ Requires-Dist: pandas
25
+ Requires-Dist: openpyxl
26
+ Requires-Dist: scikit-image
27
+ Requires-Dist: torch
28
+ Requires-Dist: torchvision
29
+ Requires-Dist: torchaudio
30
+
31
+ # SpecUNet Implementation for Spectroscopic Single-molecule Localization Microscopy(sSMLM) Imaging Denoising
32
+ This project implements a deep learning model with a U-Net-based architecture for processing single-molecule spectral localization microscopy (sSMLM) images.
33
+ The goal is to predict background signals and produce denoised spectral images that preserve useful spectral information for downstream analysis.
34
+ Implementation is highly configurable, supporting different training and testing datasets, adjustable hyperparameters, and reproducible evaluation workflows.
35
+ > Reference: Mao, H. et al. “Framework for Accurate Single-Molecule Spectroscopic Imaging Analyses Using Monte Carlo Simulation and Deep Learning,” *Analytical Chemistry* (2025). :contentReference[oaicite:4]{index=4}
36
+
37
+ ## Table of Contents
38
+ 1. [Project Structure](#project-structure)
39
+ 2. [Models](#models)
40
+ 3. [Setup](#setup)
41
+ 4. [Usage](#usage)
42
+ 5. [Contributions](#contributions)
43
+ 6. [Acknowledgments](#acknowledgments)
44
+
45
+ ## Project Structure
46
+ ```markdown
47
+ .
48
+ ├── dist/ # Stores all packaged builds (not included in repo)
49
+ ├── data/ # Stores all datasets (not included in repo)
50
+ ├── images/ # Reference images for README.md
51
+ ├── specunet_pkg/ # Main package directory
52
+ │ ├── config/ # Configuration directory
53
+ │ ├── dataset.py # Data loading and preprocessing functions
54
+ │ ├── hyperparameter_search.py# Hyperparameter tuning script
55
+ │ ├── main.py # Main entry point for training/testing
56
+ │ ├── metrics.py # Metric calculation functions
57
+ │ ├── models.py # Model definitions (UNet)
58
+ │ ├── parser.py # Parsing configuration files
59
+ │ ├── requirements.txt # List of dependencies
60
+ │ ├── test.py # Testing script for background predictions
61
+ │ ├── train.py # Training script
62
+ │ └── utils.py # Utility functions (e.g., seeds, logging)
63
+ ├── test_models/ # Saved trained models and test results (not included)
64
+ ├── .gitignore # Git ignore file
65
+ ├── pyproject.toml # Build system dependencies and config
66
+ ├── README.md # Project documentation (this file)
67
+ ├── specunet_pkg.egg-info # Stores build metadata (not included in repo)
68
+ └── setup.cfg # Package setup configuration
69
+ ```
70
+
71
+ ### SpecUNet Architecture
72
+
73
+ The SpecUNet structure takes a given image with dimensions of `16x128` and applies the following transformations:
74
+ 1. 4 Downsampling Blocks (Encoder)
75
+
76
+ Within each downsampling block consists the following:
77
+
78
+ - Convolutional 2D Layer
79
+ - BatchNorm 2D Layer
80
+ - ReLU Layer
81
+ - Max Pooling 2D Layer (kernel size of 2 and stride of 2)
82
+
83
+
84
+ 2. Bottleneck
85
+
86
+ The bottleneck for this network consists of the following:
87
+
88
+ - Convolutional 2D Layer
89
+ - ReLU Layer
90
+ - Convolutional 2D Layer
91
+ - ReLU Layer
92
+
93
+ 3. 4 Upsampling Blocks (Decoder)
94
+
95
+ Each upsampling block consists of the following:
96
+
97
+ - Convolutional Transposed 2D Layer (with kernel size of 2 and stride of 2)
98
+ - ReLU Layer
99
+ - Concatentation from nth decoder block
100
+ After each upsampling block, we add the resulting output from the nth encoder to the nth decoder.
101
+
102
+
103
+ 4. Final Convolutional Layer and ReLU Layer
104
+
105
+ We have also included the following diagram to visualize this network as well. The diagram is color-coded with
106
+ all the relevant layers, which are color-coded consistent with the legend below:
107
+
108
+ * Yellow - Convolutional Layer
109
+ * Green - BatchNorm Layer
110
+ * Purple - ReLU Layer
111
+ * Red - Pooling Layer
112
+ * Blue - Convolution Layer
113
+
114
+ ![Conventional UNet Structure](images/conventional_unet_structure.png)
115
+
116
+ ## Setup/Installation
117
+ You can install SpecUNet either by downloading a prepackaged release (recommended for general use) or by cloning the source code directly (recommended for developers and contributors).
118
+
119
+ ### Option 1: Prepackaged Release (Recommended)
120
+ If you just want to use the models and functions without modifying the underlying code, you can install the pre-built package directly from our releases.
121
+
122
+ 1. Navigate to the [Releases section](https://github.com/switfluors/SpecUNet/releases) of this repository.
123
+ 2. Download the latest .whl (wheel) file from the assets list.
124
+ 3. Open your terminal, navigate to the folder where you downloaded the file, and install it using pip:
125
+
126
+ ```Bash
127
+ pip install specunet-X.Y.Z-py3-none-any.whl
128
+ ```
129
+ (Note: Replace X.Y.Z with the actual version number you downloaded).
130
+
131
+ ### Option 2: Source Code & Local Packaging
132
+ If you want to modify the source code, run the hyperparameter search, or build the package yourself, follow these steps:
133
+
134
+ 1. Clone the repository:
135
+
136
+ ```Bash
137
+ git clone https://github.com/switfluors/SpecUNet.git
138
+ cd .\Python
139
+ ```
140
+
141
+ 2. Create a virtual environment (Optional but highly recommended):
142
+ Creating an isolated environment ensures that these dependencies don't conflict with other Python projects on your machine.
143
+
144
+ ```Bash
145
+ python -m venv venv
146
+ # On Windows:
147
+ venv\Scripts\activate
148
+ # On macOS/Linux:
149
+ source venv/bin/activate
150
+ ```
151
+
152
+ 3. Install dependencies:
153
+
154
+ ```Bash
155
+ pip install -r specunet_pkg/requirements.txt
156
+ ```
157
+ ⚠️ Note regarding PyTorch: Currently, the requirements.txt is configured for CUDA 12.8. If your machine uses a different GPU architecture or if you are running on CPU only, this may fail or run slowly. Please adjust the PyTorch installation command according to your system specs via the official PyTorch documentation.
158
+
159
+ 4. Install the package locally:
160
+ To make the specunet_pkg available anywhere on your system while keeping the code editable, install it in development mode:
161
+
162
+ ```Bash
163
+ pip install -e .
164
+ ```
165
+
166
+ 5. (optional) Packaging your own build:
167
+ If you want to build and install your own version of SpecUNet as a package, you can run the following commands:
168
+ Note: Adjust the `setup.cfg` file's version number as required.
169
+
170
+ ```Bash
171
+ pip install build
172
+ python -m build
173
+ ```
174
+
175
+ The associated pip `.whl` and `.tar.gz` files will be located in the `dist/` folder.
176
+
177
+ ### Core Dependencies
178
+ At the moment, the current Python libraries are being used:
179
+ - `h5py` and `scipy`: Importing MATLAB data files `.mat` to Python
180
+ - `torch`: Training Python models with the dataset
181
+ - `matplotlib`: General data visualization
182
+ - `numpy`: Data manipulation
183
+ - `openpyxl` and `pandas`: Storing spreadsheets of spectral metrics
184
+ - `scikit-learn` and `scipy`: Calculating various metrics
185
+ - `tqdm`: Progress monitoring
186
+ - `onnx`: For storing models in ONNX format (only works on non-5070ti GPUs)
187
+
188
+ ## Usage
189
+
190
+ ### Configuration Files
191
+
192
+ Configuration files, which are stored in the `config` directory, indicate specific setups of the dataset,
193
+ training or testing phase, model hyperparameters, dataset setups, and much more. We have included two default configurations
194
+ for both spectral-based and spatial-based datasets, which are identified in "SpecUNet.json" and "SpatUNet.json" files, respectively.
195
+
196
+ Additional configuration can be overridden when running main.py with provided flags, which are expressed in detail below.
197
+ Moreover, configuration files can be wholly ignored with main.py commands, which are discussed in later sections.
198
+
199
+ ### Loading a Dataset
200
+
201
+ This library assumes that you are using data stored in a MATLAB ".mat" format in either "MATLAB v5+" or "MATLAB v7+"
202
+ formatting, which will use either `scipy` or `h5py` libraries, respectively.
203
+
204
+ The paths to the training and testing dataset can be simply passed with the `--train_path` and `--test_path`, respectively.
205
+
206
+ ### Model Training
207
+
208
+ #### Hyperparameters
209
+
210
+ The model is trained with the following hyperparameters, which can be adjusted with the corresponding flag
211
+ defined in the `args` variable in `main.py`:
212
+
213
+ - Learning Rate (`--lr`)
214
+ - L2-regularization (`--weight_decay`)
215
+ - Batch size (`--bs`)
216
+ - Epochs (`--epochs`)
217
+ - Loss function (`--loss_fn`)
218
+ - Optimizer (`--optimizer`)
219
+
220
+ We are also using a step learning rate scheduler, which reduces the learning rate by a gamma factor
221
+ after a certain step size:
222
+
223
+ - Scheduler step size (`--scheduler_step_size`)
224
+ - Scheduler gamma (`--scheduler_gamma`)
225
+
226
+ ### Training/Testing a File
227
+
228
+ Supported flags used to train and test a model:
229
+
230
+ * `--exp_name` - Experiment name (used to save as a folder name under `test_models`)
231
+ * `--config` - Configuration file path
232
+ * `--train` - Train a model with a given train dataset
233
+ * `--test_sim` - Test a model with a given simulated dataset (target values provided)
234
+ * `--test_exp` - Test a model with a given experimental dataset (target values not provided)
235
+ * `--train_path` - Train dataset path (*.mat file)
236
+ * `--train_size` - Train dataset size
237
+ * `--test_path` - Test dataset path (*.mat file)
238
+ * `--test_size` - Test dataset size
239
+ * `--seed` - Random seed
240
+ * `--norm` - Identifies if data is normalized (for visualization purposes only)
241
+ * `--input_size` - Input shape
242
+ * `--validation_split` - Ratio to split dataset into training and validation (only applies when `--train` flag is set)
243
+ * `--num_workers` - Number of workers to load data into the GPU
244
+ * `--input_name` - Variable that stores the input (sptimg) data
245
+ * `--target_name` - Variable that stores the target (tbg) data
246
+ * `--GTspt` - Variable for ground truth spectra image data
247
+ * `--spt` - 1D spectrum profile for each spectra image (provide if available for simulated/experimental data)
248
+ * `--model_type` - Model that you want to train/test (currently, only "unet" is supported)
249
+ * `--epochs` - Number of epochs used for training
250
+ * `--bs` - Batch size
251
+ * `--lr` - Learning rate
252
+ * `--weight_decay` - L2-regularization factor
253
+ * `--loss_fn` - Loss function
254
+ * `--optimizer` - Optimization function used
255
+ * `--initializer` - Initializer used for training
256
+ * `--lr_scheduler` - Type of Learning Rate Scheduler used
257
+ * `--scheduler_step_size` - LR scheduler step size used
258
+ * `--scheduler_gamma` - LR scheduler gamma used
259
+
260
+ #### Training a Model (--train)
261
+
262
+ * If you are training a model, run the `main.py` script with the `--train` flag.
263
+ * Make sure to define the dataset and hyperparameters by calling the specific flags defined in the `args` variable in `main.py`.
264
+ * If no dataset is provided, then the training dataset will be split into training and testing by the `--train_test_split` flag.
265
+ * At the end of training, for each model set to train (defined by `--model_type` flag), it will store the following into its respective folder:
266
+ - Trained model (weights stored in .pth file, ONNX file in .onnx file)
267
+ - Training and testing RMSE losses on a per-epoch basis (stored in .npy file)
268
+ - Training and testing RMSE loss graphs (stored in .png file)
269
+ * Additionally, in the main `exp_name` directory, the following files will be provided:
270
+ - `configs` directory will store all final configurations for each image with its phase(s) and timestamp
271
+ - Logging file (stored in .log file)
272
+
273
+ #### Testing a Model (--test_sim or --test_exp)
274
+ If you are testing a model along with training, or simply testing an existing trained model, just add the `--test_sim`
275
+ flag for simulated data (where ground truth variables are present), or `--test_exp` flag for experimental data (where
276
+ ground truth variables are not present). This will produce several figures for evaluation purposes, including:
277
+
278
+ 1. An RMSE histogram to compare the performance of all trained UNet models within a specific test folder (only with `--test_sim`)
279
+ 2. For each model, it will also store the following, which will showcase:
280
+ - Up to five representative images, which include:
281
+ - Predicted Background (`--test_sim` only)
282
+ - Ground Truth Background
283
+ - Predicted Spectra (`--test_sim` only)
284
+ - Ground Truth Spectra
285
+ - Original Spectra
286
+ - Can optionally store all representative images for each model as TIFF files
287
+
288
+ Configuration files will define the overall structure and output of each train or test phase, and can be modified
289
+ accordingly.
290
+
291
+ Note: If you are training and testing a model at the same time with two different datasets, make sure that the
292
+ input, target, and ground spectral data (if applicable) use consistent variable names within both datasets.
293
+
294
+ An example of running a model to train and test is shown below.
295
+
296
+ #### Example
297
+ As an example, we would like to train SpecUNet on the `Sample_TrainingData_10000.mat` dataset and test on the simulated `TestingData.mat` dataset,
298
+ target variable `tbg4`, and ground truth spectral image variable `GTspt` that are stored in the `data` folder with a learning rate of 0.001,
299
+ L2-regularization factor of 0.0005, and loss function of MAE.
300
+ Moreover, you would like to store the trained model with an experiment name of `Trained_Model1`.
301
+
302
+ In this case, the command will be as simple as:
303
+
304
+ ```bash
305
+ python -m specunet_pkg --train --test_sim --exp_name "Trained_Model1" --model_type "unet" --train_path "data/Sample_TrainingData_10000.mat" --train_size 10000 --test_path "data/TestingData.mat" --test_size 5000 --input_name "sptimg4" --target_name "tbg4" --GTspt "GTspt" --lr 0.001 --weight_decay 0.0005 --loss_fn "mae"
306
+ ```
307
+
308
+ Alternatively, you can use a configuration file to define the same training and testing parameters, as long as it is defined by your expectations:
309
+
310
+ ```bash
311
+ python -m specunet_pkg --train --test_sim --exp_name "Trained_Model1" --config "config/SpecUNet.json"
312
+ ```
313
+
314
+ If you would like to only test selected model under "Trained_Model1" on a specific dataset you would run the following command:
315
+
316
+ ## If Predicting Simulated Data with ground truth
317
+ ```bash
318
+ python -m specunet_pkg --test_sim --exp_name "Trained_Model1" --model_type unet --test_path "data/TestingData.mat" --test_size 5000 --input_name "sptimg4" --target_name "tbg4" --GTspt "GTspt"
319
+ ```
320
+ Alternatively use config data:
321
+ ```bash
322
+ python main.py --test_sim --exp_name "Trained_Model1" --config "config/SpecUNet.json"
323
+ ```
324
+
325
+ If you would like to then to test the pretrained model under "Trained_Model1" on an experimental dataset called `ExpTestingData.mat` with 1163
326
+ samples and target variable of `final_bbimg`, you would run the following command:
327
+
328
+ ## If Predicting Experimental Data without ground truth
329
+ ```bash
330
+ python -m specunet_pkg --test_exp --exp_name "Trained_Model1" --test_path "Data/ExpTestingData.mat" --test_size 1163 --input_name "final_bbimg" --model_type "unet"
331
+ ```
332
+
333
+ Alternatively use config data:
334
+ ```bash
335
+ python main.py --test_exp --exp_name "Trained_Model1" --config "config/SpecUNet.json"
336
+ ```
337
+
338
+ ### Hyperparameter Tuning
339
+
340
+ The `hyperparameter_search.py` file allows for optimizing the performance of a given model on a dataset by
341
+ testing the search space of the given hyperparameter that can be called via `python main.py --train`, similar to
342
+ GridSearchCV class in `scikit-learn`, though not as efficient.
343
+
344
+ ## Acknowledgments
345
+
346
+ We thank Dr. Dongkuan Xu and Dr. Caroline Laplante for their guidance on this project.
@@ -0,0 +1,316 @@
1
+ # SpecUNet Implementation for Spectroscopic Single-molecule Localization Microscopy(sSMLM) Imaging Denoising
2
+ This project implements a deep learning model with a U-Net-based architecture for processing single-molecule spectral localization microscopy (sSMLM) images.
3
+ The goal is to predict background signals and produce denoised spectral images that preserve useful spectral information for downstream analysis.
4
+ Implementation is highly configurable, supporting different training and testing datasets, adjustable hyperparameters, and reproducible evaluation workflows.
5
+ > Reference: Mao, H. et al. “Framework for Accurate Single-Molecule Spectroscopic Imaging Analyses Using Monte Carlo Simulation and Deep Learning,” *Analytical Chemistry* (2025). :contentReference[oaicite:4]{index=4}
6
+
7
+ ## Table of Contents
8
+ 1. [Project Structure](#project-structure)
9
+ 2. [Models](#models)
10
+ 3. [Setup](#setup)
11
+ 4. [Usage](#usage)
12
+ 5. [Contributions](#contributions)
13
+ 6. [Acknowledgments](#acknowledgments)
14
+
15
+ ## Project Structure
16
+ ```markdown
17
+ .
18
+ ├── dist/ # Stores all packaged builds (not included in repo)
19
+ ├── data/ # Stores all datasets (not included in repo)
20
+ ├── images/ # Reference images for README.md
21
+ ├── specunet_pkg/ # Main package directory
22
+ │ ├── config/ # Configuration directory
23
+ │ ├── dataset.py # Data loading and preprocessing functions
24
+ │ ├── hyperparameter_search.py# Hyperparameter tuning script
25
+ │ ├── main.py # Main entry point for training/testing
26
+ │ ├── metrics.py # Metric calculation functions
27
+ │ ├── models.py # Model definitions (UNet)
28
+ │ ├── parser.py # Parsing configuration files
29
+ │ ├── requirements.txt # List of dependencies
30
+ │ ├── test.py # Testing script for background predictions
31
+ │ ├── train.py # Training script
32
+ │ └── utils.py # Utility functions (e.g., seeds, logging)
33
+ ├── test_models/ # Saved trained models and test results (not included)
34
+ ├── .gitignore # Git ignore file
35
+ ├── pyproject.toml # Build system dependencies and config
36
+ ├── README.md # Project documentation (this file)
37
+ ├── specunet_pkg.egg-info # Stores build metadata (not included in repo)
38
+ └── setup.cfg # Package setup configuration
39
+ ```
40
+
41
+ ### SpecUNet Architecture
42
+
43
+ The SpecUNet structure takes a given image with dimensions of `16x128` and applies the following transformations:
44
+ 1. 4 Downsampling Blocks (Encoder)
45
+
46
+ Within each downsampling block consists the following:
47
+
48
+ - Convolutional 2D Layer
49
+ - BatchNorm 2D Layer
50
+ - ReLU Layer
51
+ - Max Pooling 2D Layer (kernel size of 2 and stride of 2)
52
+
53
+
54
+ 2. Bottleneck
55
+
56
+ The bottleneck for this network consists of the following:
57
+
58
+ - Convolutional 2D Layer
59
+ - ReLU Layer
60
+ - Convolutional 2D Layer
61
+ - ReLU Layer
62
+
63
+ 3. 4 Upsampling Blocks (Decoder)
64
+
65
+ Each upsampling block consists of the following:
66
+
67
+ - Convolutional Transposed 2D Layer (with kernel size of 2 and stride of 2)
68
+ - ReLU Layer
69
+ - Concatentation from nth decoder block
70
+ After each upsampling block, we add the resulting output from the nth encoder to the nth decoder.
71
+
72
+
73
+ 4. Final Convolutional Layer and ReLU Layer
74
+
75
+ We have also included the following diagram to visualize this network as well. The diagram is color-coded with
76
+ all the relevant layers, which are color-coded consistent with the legend below:
77
+
78
+ * Yellow - Convolutional Layer
79
+ * Green - BatchNorm Layer
80
+ * Purple - ReLU Layer
81
+ * Red - Pooling Layer
82
+ * Blue - Convolution Layer
83
+
84
+ ![Conventional UNet Structure](images/conventional_unet_structure.png)
85
+
86
+ ## Setup/Installation
87
+ You can install SpecUNet either by downloading a prepackaged release (recommended for general use) or by cloning the source code directly (recommended for developers and contributors).
88
+
89
+ ### Option 1: Prepackaged Release (Recommended)
90
+ If you just want to use the models and functions without modifying the underlying code, you can install the pre-built package directly from our releases.
91
+
92
+ 1. Navigate to the [Releases section](https://github.com/switfluors/SpecUNet/releases) of this repository.
93
+ 2. Download the latest .whl (wheel) file from the assets list.
94
+ 3. Open your terminal, navigate to the folder where you downloaded the file, and install it using pip:
95
+
96
+ ```Bash
97
+ pip install specunet-X.Y.Z-py3-none-any.whl
98
+ ```
99
+ (Note: Replace X.Y.Z with the actual version number you downloaded).
100
+
101
+ ### Option 2: Source Code & Local Packaging
102
+ If you want to modify the source code, run the hyperparameter search, or build the package yourself, follow these steps:
103
+
104
+ 1. Clone the repository:
105
+
106
+ ```Bash
107
+ git clone https://github.com/switfluors/SpecUNet.git
108
+ cd .\Python
109
+ ```
110
+
111
+ 2. Create a virtual environment (Optional but highly recommended):
112
+ Creating an isolated environment ensures that these dependencies don't conflict with other Python projects on your machine.
113
+
114
+ ```Bash
115
+ python -m venv venv
116
+ # On Windows:
117
+ venv\Scripts\activate
118
+ # On macOS/Linux:
119
+ source venv/bin/activate
120
+ ```
121
+
122
+ 3. Install dependencies:
123
+
124
+ ```Bash
125
+ pip install -r specunet_pkg/requirements.txt
126
+ ```
127
+ ⚠️ Note regarding PyTorch: Currently, the requirements.txt is configured for CUDA 12.8. If your machine uses a different GPU architecture or if you are running on CPU only, this may fail or run slowly. Please adjust the PyTorch installation command according to your system specs via the official PyTorch documentation.
128
+
129
+ 4. Install the package locally:
130
+ To make the specunet_pkg available anywhere on your system while keeping the code editable, install it in development mode:
131
+
132
+ ```Bash
133
+ pip install -e .
134
+ ```
135
+
136
+ 5. (optional) Packaging your own build:
137
+ If you want to build and install your own version of SpecUNet as a package, you can run the following commands:
138
+ Note: Adjust the `setup.cfg` file's version number as required.
139
+
140
+ ```Bash
141
+ pip install build
142
+ python -m build
143
+ ```
144
+
145
+ The associated pip `.whl` and `.tar.gz` files will be located in the `dist/` folder.
146
+
147
+ ### Core Dependencies
148
+ At the moment, the current Python libraries are being used:
149
+ - `h5py` and `scipy`: Importing MATLAB data files `.mat` to Python
150
+ - `torch`: Training Python models with the dataset
151
+ - `matplotlib`: General data visualization
152
+ - `numpy`: Data manipulation
153
+ - `openpyxl` and `pandas`: Storing spreadsheets of spectral metrics
154
+ - `scikit-learn` and `scipy`: Calculating various metrics
155
+ - `tqdm`: Progress monitoring
156
+ - `onnx`: For storing models in ONNX format (only works on non-5070ti GPUs)
157
+
158
+ ## Usage
159
+
160
+ ### Configuration Files
161
+
162
+ Configuration files, which are stored in the `config` directory, indicate specific setups of the dataset,
163
+ training or testing phase, model hyperparameters, dataset setups, and much more. We have included two default configurations
164
+ for both spectral-based and spatial-based datasets, which are identified in "SpecUNet.json" and "SpatUNet.json" files, respectively.
165
+
166
+ Additional configuration can be overridden when running main.py with provided flags, which are expressed in detail below.
167
+ Moreover, configuration files can be wholly ignored with main.py commands, which are discussed in later sections.
168
+
169
+ ### Loading a Dataset
170
+
171
+ This library assumes that you are using data stored in a MATLAB ".mat" format in either "MATLAB v5+" or "MATLAB v7+"
172
+ formatting, which will use either `scipy` or `h5py` libraries, respectively.
173
+
174
+ The paths to the training and testing dataset can be simply passed with the `--train_path` and `--test_path`, respectively.
175
+
176
+ ### Model Training
177
+
178
+ #### Hyperparameters
179
+
180
+ The model is trained with the following hyperparameters, which can be adjusted with the corresponding flag
181
+ defined in the `args` variable in `main.py`:
182
+
183
+ - Learning Rate (`--lr`)
184
+ - L2-regularization (`--weight_decay`)
185
+ - Batch size (`--bs`)
186
+ - Epochs (`--epochs`)
187
+ - Loss function (`--loss_fn`)
188
+ - Optimizer (`--optimizer`)
189
+
190
+ We are also using a step learning rate scheduler, which reduces the learning rate by a gamma factor
191
+ after a certain step size:
192
+
193
+ - Scheduler step size (`--scheduler_step_size`)
194
+ - Scheduler gamma (`--scheduler_gamma`)
195
+
196
+ ### Training/Testing a File
197
+
198
+ Supported flags used to train and test a model:
199
+
200
+ * `--exp_name` - Experiment name (used to save as a folder name under `test_models`)
201
+ * `--config` - Configuration file path
202
+ * `--train` - Train a model with a given train dataset
203
+ * `--test_sim` - Test a model with a given simulated dataset (target values provided)
204
+ * `--test_exp` - Test a model with a given experimental dataset (target values not provided)
205
+ * `--train_path` - Train dataset path (*.mat file)
206
+ * `--train_size` - Train dataset size
207
+ * `--test_path` - Test dataset path (*.mat file)
208
+ * `--test_size` - Test dataset size
209
+ * `--seed` - Random seed
210
+ * `--norm` - Identifies if data is normalized (for visualization purposes only)
211
+ * `--input_size` - Input shape
212
+ * `--validation_split` - Ratio to split dataset into training and validation (only applies when `--train` flag is set)
213
+ * `--num_workers` - Number of workers to load data into the GPU
214
+ * `--input_name` - Variable that stores the input (sptimg) data
215
+ * `--target_name` - Variable that stores the target (tbg) data
216
+ * `--GTspt` - Variable for ground truth spectra image data
217
+ * `--spt` - 1D spectrum profile for each spectra image (provide if available for simulated/experimental data)
218
+ * `--model_type` - Model that you want to train/test (currently, only "unet" is supported)
219
+ * `--epochs` - Number of epochs used for training
220
+ * `--bs` - Batch size
221
+ * `--lr` - Learning rate
222
+ * `--weight_decay` - L2-regularization factor
223
+ * `--loss_fn` - Loss function
224
+ * `--optimizer` - Optimization function used
225
+ * `--initializer` - Initializer used for training
226
+ * `--lr_scheduler` - Type of Learning Rate Scheduler used
227
+ * `--scheduler_step_size` - LR scheduler step size used
228
+ * `--scheduler_gamma` - LR scheduler gamma used
229
+
230
+ #### Training a Model (--train)
231
+
232
+ * If you are training a model, run the `main.py` script with the `--train` flag.
233
+ * Make sure to define the dataset and hyperparameters by calling the specific flags defined in the `args` variable in `main.py`.
234
+ * If no dataset is provided, then the training dataset will be split into training and testing by the `--train_test_split` flag.
235
+ * At the end of training, for each model set to train (defined by `--model_type` flag), it will store the following into its respective folder:
236
+ - Trained model (weights stored in .pth file, ONNX file in .onnx file)
237
+ - Training and testing RMSE losses on a per-epoch basis (stored in .npy file)
238
+ - Training and testing RMSE loss graphs (stored in .png file)
239
+ * Additionally, in the main `exp_name` directory, the following files will be provided:
240
+ - `configs` directory will store all final configurations for each image with its phase(s) and timestamp
241
+ - Logging file (stored in .log file)
242
+
243
+ #### Testing a Model (--test_sim or --test_exp)
244
+ If you are testing a model along with training, or simply testing an existing trained model, just add the `--test_sim`
245
+ flag for simulated data (where ground truth variables are present), or `--test_exp` flag for experimental data (where
246
+ ground truth variables are not present). This will produce several figures for evaluation purposes, including:
247
+
248
+ 1. An RMSE histogram to compare the performance of all trained UNet models within a specific test folder (only with `--test_sim`)
249
+ 2. For each model, it will also store the following, which will showcase:
250
+ - Up to five representative images, which include:
251
+ - Predicted Background (`--test_sim` only)
252
+ - Ground Truth Background
253
+ - Predicted Spectra (`--test_sim` only)
254
+ - Ground Truth Spectra
255
+ - Original Spectra
256
+ - Can optionally store all representative images for each model as TIFF files
257
+
258
+ Configuration files will define the overall structure and output of each train or test phase, and can be modified
259
+ accordingly.
260
+
261
+ Note: If you are training and testing a model at the same time with two different datasets, make sure that the
262
+ input, target, and ground spectral data (if applicable) use consistent variable names within both datasets.
263
+
264
+ An example of running a model to train and test is shown below.
265
+
266
+ #### Example
267
+ As an example, we would like to train SpecUNet on the `Sample_TrainingData_10000.mat` dataset and test on the simulated `TestingData.mat` dataset,
268
+ target variable `tbg4`, and ground truth spectral image variable `GTspt` that are stored in the `data` folder with a learning rate of 0.001,
269
+ L2-regularization factor of 0.0005, and loss function of MAE.
270
+ Moreover, you would like to store the trained model with an experiment name of `Trained_Model1`.
271
+
272
+ In this case, the command will be as simple as:
273
+
274
+ ```bash
275
+ python -m specunet_pkg --train --test_sim --exp_name "Trained_Model1" --model_type "unet" --train_path "data/Sample_TrainingData_10000.mat" --train_size 10000 --test_path "data/TestingData.mat" --test_size 5000 --input_name "sptimg4" --target_name "tbg4" --GTspt "GTspt" --lr 0.001 --weight_decay 0.0005 --loss_fn "mae"
276
+ ```
277
+
278
+ Alternatively, you can use a configuration file to define the same training and testing parameters, as long as it is defined by your expectations:
279
+
280
+ ```bash
281
+ python -m specunet_pkg --train --test_sim --exp_name "Trained_Model1" --config "config/SpecUNet.json"
282
+ ```
283
+
284
+ If you would like to only test selected model under "Trained_Model1" on a specific dataset you would run the following command:
285
+
286
+ ## If Predicting Simulated Data with ground truth
287
+ ```bash
288
+ python -m specunet_pkg --test_sim --exp_name "Trained_Model1" --model_type unet --test_path "data/TestingData.mat" --test_size 5000 --input_name "sptimg4" --target_name "tbg4" --GTspt "GTspt"
289
+ ```
290
+ Alternatively use config data:
291
+ ```bash
292
+ python main.py --test_sim --exp_name "Trained_Model1" --config "config/SpecUNet.json"
293
+ ```
294
+
295
+ If you would like to then to test the pretrained model under "Trained_Model1" on an experimental dataset called `ExpTestingData.mat` with 1163
296
+ samples and target variable of `final_bbimg`, you would run the following command:
297
+
298
+ ## If Predicting Experimental Data without ground truth
299
+ ```bash
300
+ python -m specunet_pkg --test_exp --exp_name "Trained_Model1" --test_path "Data/ExpTestingData.mat" --test_size 1163 --input_name "final_bbimg" --model_type "unet"
301
+ ```
302
+
303
+ Alternatively use config data:
304
+ ```bash
305
+ python main.py --test_exp --exp_name "Trained_Model1" --config "config/SpecUNet.json"
306
+ ```
307
+
308
+ ### Hyperparameter Tuning
309
+
310
+ The `hyperparameter_search.py` file allows for optimizing the performance of a given model on a dataset by
311
+ testing the search space of the given hyperparameter that can be called via `python main.py --train`, similar to
312
+ GridSearchCV class in `scikit-learn`, though not as efficient.
313
+
314
+ ## Acknowledgments
315
+
316
+ We thank Dr. Dongkuan Xu and Dr. Caroline Laplante for their guidance on this project.
@@ -0,0 +1,3 @@
1
+ [build-system]
2
+ requires = ["setuptools>=61.0"]
3
+ build-backend = "setuptools.build_meta"