viseda 1.0.0__tar.gz → 1.0.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- viseda-1.0.1/PKG-INFO +731 -0
- viseda-1.0.1/README.md +687 -0
- {viseda-1.0.0 → viseda-1.0.1}/pyproject.toml +6 -6
- {viseda-1.0.0 → viseda-1.0.1}/viseda/__init__.py +1 -1
- viseda-1.0.1/viseda.egg-info/PKG-INFO +731 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda.egg-info/requires.txt +5 -5
- viseda-1.0.0/PKG-INFO +0 -266
- viseda-1.0.0/README.md +0 -222
- viseda-1.0.0/viseda.egg-info/PKG-INFO +0 -266
- {viseda-1.0.0 → viseda-1.0.1}/LICENSE +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/setup.cfg +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/tests/test_smoke.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/cli.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/core/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/core/base.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/hyperspectral/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/hyperspectral/eda.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/image/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/image/eda.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/pointcloud/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/pointcloud/eda.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/report/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/report/html_report.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/text/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/text/eda.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/utils/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/utils/helpers.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/video/__init__.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda/video/eda.py +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda.egg-info/SOURCES.txt +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda.egg-info/dependency_links.txt +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda.egg-info/entry_points.txt +0 -0
- {viseda-1.0.0 → viseda-1.0.1}/viseda.egg-info/top_level.txt +0 -0
viseda-1.0.1/PKG-INFO
ADDED
|
@@ -0,0 +1,731 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: viseda
|
|
3
|
+
Version: 1.0.1
|
|
4
|
+
Summary: Visual Exploratory Data Analysis for image, video, hyperspectral, point-cloud, and text/NLP datasets
|
|
5
|
+
Author: Isaac Osei Agyemang, Daniel Acheampong, Adu Asare Baffour, Isaac Adjei-Mensah
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Keywords: EDA,exploratory data analysis,image,video,hyperspectral,point cloud,LiDAR,text,NLP,computer vision,remote sensing
|
|
8
|
+
Classifier: Development Status :: 4 - Beta
|
|
9
|
+
Classifier: Intended Audience :: Science/Research
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: Programming Language :: Python :: 3
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
16
|
+
Classifier: Topic :: Scientific/Engineering :: Image Processing
|
|
17
|
+
Classifier: Topic :: Scientific/Engineering :: Visualization
|
|
18
|
+
Classifier: Topic :: Scientific/Engineering :: GIS
|
|
19
|
+
Requires-Python: >=3.9
|
|
20
|
+
Description-Content-Type: text/markdown
|
|
21
|
+
License-File: LICENSE
|
|
22
|
+
Requires-Dist: numpy<2,>=1.23
|
|
23
|
+
Requires-Dist: matplotlib<4,>=3.6
|
|
24
|
+
Requires-Dist: scikit-learn<2,>=1.2
|
|
25
|
+
Requires-Dist: scikit-image<1,>=0.20
|
|
26
|
+
Requires-Dist: opencv-python<5,>=4.7
|
|
27
|
+
Provides-Extra: hyperspectral
|
|
28
|
+
Requires-Dist: scipy>=1.10; extra == "hyperspectral"
|
|
29
|
+
Requires-Dist: spectral>=0.23; extra == "hyperspectral"
|
|
30
|
+
Requires-Dist: rasterio>=1.3; extra == "hyperspectral"
|
|
31
|
+
Provides-Extra: pointcloud
|
|
32
|
+
Requires-Dist: laspy[lazrs]>=2.3; extra == "pointcloud"
|
|
33
|
+
Provides-Extra: all
|
|
34
|
+
Requires-Dist: scipy>=1.10; extra == "all"
|
|
35
|
+
Requires-Dist: spectral>=0.23; extra == "all"
|
|
36
|
+
Requires-Dist: rasterio>=1.3; extra == "all"
|
|
37
|
+
Requires-Dist: laspy[lazrs]>=2.3; extra == "all"
|
|
38
|
+
Provides-Extra: dev
|
|
39
|
+
Requires-Dist: pytest>=7.4; extra == "dev"
|
|
40
|
+
Requires-Dist: pytest-cov>=4.1; extra == "dev"
|
|
41
|
+
Requires-Dist: build>=1.0; extra == "dev"
|
|
42
|
+
Requires-Dist: twine>=5.0; extra == "dev"
|
|
43
|
+
Dynamic: license-file
|
|
44
|
+
|
|
45
|
+
<div align="center">
|
|
46
|
+
|
|
47
|
+
<img src="VisEDA.png" alt="VisEDA — Comprehensive Documentation" width="100%">
|
|
48
|
+
|
|
49
|
+
# VisEDA
|
|
50
|
+
|
|
51
|
+
### Visual Exploratory Data Analysis for multimodal datasets
|
|
52
|
+
|
|
53
|
+
**ImageEDA · VideoEDA · HyperspectralEDA · PointCloudEDA · TextEDA**
|
|
54
|
+
|
|
55
|
+
[](https://pypi.org/project/viseda/)
|
|
56
|
+
[](https://pypi.org/project/viseda/)
|
|
57
|
+
[](LICENSE)
|
|
58
|
+
|
|
59
|
+
**One toolkit for inspecting, understanding, and validating visual and text datasets before modelling.**
|
|
60
|
+
|
|
61
|
+
</div>
|
|
62
|
+
|
|
63
|
+
---
|
|
64
|
+
|
|
65
|
+
## Overview
|
|
66
|
+
|
|
67
|
+
**VisEDA** is a Python toolkit for exploratory data analysis across five major data modalities:
|
|
68
|
+
|
|
69
|
+
| Module | Designed for | Core analyses |
|
|
70
|
+
|---|---|---|
|
|
71
|
+
| **ImageEDA** | Image datasets | size, aspect ratio, brightness, contrast, sharpness, noise, exposure, colour, texture, frequency, duplicates, class balance |
|
|
72
|
+
| **VideoEDA** | Video datasets | frame count, duration, FPS, temporal brightness/contrast, sharpness, blur, motion, scene changes, similarity |
|
|
73
|
+
| **HyperspectralEDA** | Hyperspectral cubes | band statistics, SNR, noise, dropout bands, vegetation/water indices, PCA, false colour, spectral diversity |
|
|
74
|
+
| **PointCloudEDA** | 3D point clouds | geometry, density, height, duplicates, outliers, nearest neighbours, PCA shape descriptors, cloud distances |
|
|
75
|
+
| **TextEDA** | Text/NLP datasets | document length, vocabulary, lexical diversity, readability, symbols, scripts, duplicates, TF-IDF distances, label analysis |
|
|
76
|
+
|
|
77
|
+
VisEDA is intended for **data scientists, researchers, students, and practitioners** who want a consistent EDA workflow before training machine-learning or deep-learning models.
|
|
78
|
+
|
|
79
|
+
---
|
|
80
|
+
|
|
81
|
+
## Why VisEDA?
|
|
82
|
+
|
|
83
|
+
Exploratory analysis is standard practice for tabular data, but visual and multimodal datasets often require separate tools and custom scripts.
|
|
84
|
+
|
|
85
|
+
VisEDA provides a unified interface for:
|
|
86
|
+
|
|
87
|
+
- **dataset inventory and integrity checks**
|
|
88
|
+
- **quality and distribution analysis**
|
|
89
|
+
- **class-balance inspection**
|
|
90
|
+
- **duplicate and similarity analysis**
|
|
91
|
+
- **per-sample and dataset-level visualisations**
|
|
92
|
+
- **self-contained HTML reports**
|
|
93
|
+
- **command-line workflows**
|
|
94
|
+
- **Python API workflows**
|
|
95
|
+
- **reproducible pre-model data auditing**
|
|
96
|
+
|
|
97
|
+
The goal is simple:
|
|
98
|
+
|
|
99
|
+
> **Understand the dataset before trusting the model.**
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## Installation
|
|
104
|
+
|
|
105
|
+
Install the base package from PyPI:
|
|
106
|
+
|
|
107
|
+
```bash
|
|
108
|
+
pip install viseda
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
### Optional extras
|
|
112
|
+
|
|
113
|
+
For extended hyperspectral file support:
|
|
114
|
+
|
|
115
|
+
```bash
|
|
116
|
+
pip install "viseda[hyperspectral]"
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
For LAS/LAZ point-cloud support:
|
|
120
|
+
|
|
121
|
+
```bash
|
|
122
|
+
pip install "viseda[pointcloud]"
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Install all optional dependencies:
|
|
126
|
+
|
|
127
|
+
```bash
|
|
128
|
+
pip install "viseda[all]"
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
Python **3.9+** is recommended.
|
|
132
|
+
|
|
133
|
+
---
|
|
134
|
+
|
|
135
|
+
# Quick Start
|
|
136
|
+
|
|
137
|
+
|
|
138
|
+
> **Important:** ImageEDA uses `plot()` for its comprehensive dataset dashboard.
|
|
139
|
+
> VideoEDA, HyperspectralEDA, PointCloudEDA, and TextEDA expose `plot_dataset()` for dataset-level dashboards.
|
|
140
|
+
|
|
141
|
+
## ImageEDA
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
from viseda import ImageEDA
|
|
145
|
+
|
|
146
|
+
eda = ImageEDA(verbose=True)
|
|
147
|
+
|
|
148
|
+
eda.load(
|
|
149
|
+
"path/to/images",
|
|
150
|
+
label_from_parent=True
|
|
151
|
+
)
|
|
152
|
+
|
|
153
|
+
summary = eda.summary()
|
|
154
|
+
|
|
155
|
+
print(summary["inventory"])
|
|
156
|
+
print(summary["quality"])
|
|
157
|
+
|
|
158
|
+
# ImageEDA's full dataset dashboard is plot(), not plot_dataset()
|
|
159
|
+
eda.plot(
|
|
160
|
+
save_path="image_dashboard.png"
|
|
161
|
+
)
|
|
162
|
+
|
|
163
|
+
eda.report(
|
|
164
|
+
"viseda_image_report.html"
|
|
165
|
+
)
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
For in-memory images:
|
|
169
|
+
|
|
170
|
+
```python
|
|
171
|
+
import numpy as np
|
|
172
|
+
from viseda import ImageEDA
|
|
173
|
+
|
|
174
|
+
images = [
|
|
175
|
+
np.random.default_rng(0).integers(
|
|
176
|
+
0, 256, size=(128, 128, 3), dtype=np.uint8
|
|
177
|
+
),
|
|
178
|
+
np.random.default_rng(1).integers(
|
|
179
|
+
0, 256, size=(128, 128, 3), dtype=np.uint8
|
|
180
|
+
),
|
|
181
|
+
]
|
|
182
|
+
|
|
183
|
+
eda = ImageEDA(
|
|
184
|
+
verbose=False,
|
|
185
|
+
compute_glcm=False,
|
|
186
|
+
compute_freq=False,
|
|
187
|
+
)
|
|
188
|
+
|
|
189
|
+
eda.load_arrays(
|
|
190
|
+
images,
|
|
191
|
+
labels=["class_a", "class_b"]
|
|
192
|
+
)
|
|
193
|
+
|
|
194
|
+
print(eda.summary()["inventory"])
|
|
195
|
+
|
|
196
|
+
eda.plot(
|
|
197
|
+
save_path="image_dashboard.png"
|
|
198
|
+
)
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
## VideoEDA
|
|
204
|
+
|
|
205
|
+
```python
|
|
206
|
+
from viseda import VideoEDA
|
|
207
|
+
|
|
208
|
+
eda = VideoEDA(
|
|
209
|
+
verbose=True,
|
|
210
|
+
frame_sample_rate=5
|
|
211
|
+
)
|
|
212
|
+
|
|
213
|
+
eda.load(
|
|
214
|
+
"path/to/videos",
|
|
215
|
+
label_from_parent=True
|
|
216
|
+
)
|
|
217
|
+
|
|
218
|
+
summary = eda.summary()
|
|
219
|
+
|
|
220
|
+
print(summary["inventory"])
|
|
221
|
+
print(summary["temporal"])
|
|
222
|
+
print(summary["motion"])
|
|
223
|
+
|
|
224
|
+
eda.plot_dataset(
|
|
225
|
+
save_path="video_dashboard.png"
|
|
226
|
+
)
|
|
227
|
+
|
|
228
|
+
eda.report(
|
|
229
|
+
"viseda_video_report.html"
|
|
230
|
+
)
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
For one video rather than the whole dataset:
|
|
234
|
+
|
|
235
|
+
```python
|
|
236
|
+
eda.plot(
|
|
237
|
+
video_index=0,
|
|
238
|
+
save_path="single_video.png"
|
|
239
|
+
)
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
---
|
|
243
|
+
|
|
244
|
+
## HyperspectralEDA
|
|
245
|
+
|
|
246
|
+
```python
|
|
247
|
+
import numpy as np
|
|
248
|
+
from viseda import HyperspectralEDA
|
|
249
|
+
|
|
250
|
+
# One wavelength value per spectral band
|
|
251
|
+
wavelengths = np.load(
|
|
252
|
+
"path/to/wavelengths.npy"
|
|
253
|
+
)
|
|
254
|
+
|
|
255
|
+
eda = HyperspectralEDA(
|
|
256
|
+
wavelengths=wavelengths,
|
|
257
|
+
compute_glcm=False,
|
|
258
|
+
compute_pca=True,
|
|
259
|
+
)
|
|
260
|
+
|
|
261
|
+
# File-backed loading is recommended when using compute_index()
|
|
262
|
+
# or pca_scores().
|
|
263
|
+
eda.load(
|
|
264
|
+
"path/to/cube.npy"
|
|
265
|
+
)
|
|
266
|
+
|
|
267
|
+
summary = eda.summary()
|
|
268
|
+
|
|
269
|
+
print(summary["inventory"])
|
|
270
|
+
print(summary["spectral_quality"])
|
|
271
|
+
|
|
272
|
+
ndvi = eda.compute_index(
|
|
273
|
+
cube_index=0,
|
|
274
|
+
index_name="ndvi"
|
|
275
|
+
)
|
|
276
|
+
|
|
277
|
+
scores, variance_ratio = eda.pca_scores(
|
|
278
|
+
cube_index=0,
|
|
279
|
+
n_components=3
|
|
280
|
+
)
|
|
281
|
+
|
|
282
|
+
eda.plot(
|
|
283
|
+
cube_index=0,
|
|
284
|
+
save_path="hyperspectral_cube.png"
|
|
285
|
+
)
|
|
286
|
+
|
|
287
|
+
eda.report(
|
|
288
|
+
"viseda_hyper_report.html"
|
|
289
|
+
)
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
For a directory of hyperspectral cubes:
|
|
293
|
+
|
|
294
|
+
```python
|
|
295
|
+
from viseda import HyperspectralEDA
|
|
296
|
+
|
|
297
|
+
eda = HyperspectralEDA()
|
|
298
|
+
|
|
299
|
+
eda.load(
|
|
300
|
+
"path/to/hyperspectral_dataset",
|
|
301
|
+
label_from_parent=True
|
|
302
|
+
)
|
|
303
|
+
|
|
304
|
+
eda.plot_dataset(
|
|
305
|
+
save_path="hyperspectral_dataset.png"
|
|
306
|
+
)
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
> **Current 1.0.0 note:** `load_arrays()` is suitable for summary and plotting workflows, but `compute_index()` and `pca_scores()` reload the selected cube from its source path. Use `load()` with a file-backed cube when calling those two methods.
|
|
310
|
+
|
|
311
|
+
---
|
|
312
|
+
|
|
313
|
+
## PointCloudEDA
|
|
314
|
+
|
|
315
|
+
```python
|
|
316
|
+
import numpy as np
|
|
317
|
+
from viseda import PointCloudEDA
|
|
318
|
+
|
|
319
|
+
points = np.random.default_rng(0).random(
|
|
320
|
+
(10000, 3)
|
|
321
|
+
).astype("float32")
|
|
322
|
+
|
|
323
|
+
eda = PointCloudEDA(
|
|
324
|
+
max_points_per_cloud=200000,
|
|
325
|
+
compute_neighbors=True,
|
|
326
|
+
compute_geometry=True,
|
|
327
|
+
)
|
|
328
|
+
|
|
329
|
+
eda.load_arrays(
|
|
330
|
+
[points],
|
|
331
|
+
labels=["sample"]
|
|
332
|
+
)
|
|
333
|
+
|
|
334
|
+
summary = eda.summary()
|
|
335
|
+
|
|
336
|
+
print(summary["inventory"])
|
|
337
|
+
print(summary["geometry"])
|
|
338
|
+
print(summary["quality"])
|
|
339
|
+
|
|
340
|
+
eda.plot_dataset(
|
|
341
|
+
save_path="pointcloud_dashboard.png"
|
|
342
|
+
)
|
|
343
|
+
|
|
344
|
+
eda.report(
|
|
345
|
+
"viseda_pointcloud_report.html"
|
|
346
|
+
)
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
For an individual cloud:
|
|
350
|
+
|
|
351
|
+
```python
|
|
352
|
+
eda.plot(
|
|
353
|
+
cloud_index=0,
|
|
354
|
+
save_path="single_cloud.png"
|
|
355
|
+
)
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
---
|
|
359
|
+
|
|
360
|
+
## TextEDA
|
|
361
|
+
|
|
362
|
+
```python
|
|
363
|
+
from viseda import TextEDA
|
|
364
|
+
|
|
365
|
+
eda = TextEDA()
|
|
366
|
+
|
|
367
|
+
eda.load_texts(
|
|
368
|
+
[
|
|
369
|
+
"Exploratory data analysis should come before model training.",
|
|
370
|
+
"TextEDA analyses vocabulary, length, readability and duplicates.",
|
|
371
|
+
],
|
|
372
|
+
labels=["eda", "eda"],
|
|
373
|
+
)
|
|
374
|
+
|
|
375
|
+
summary = eda.summary()
|
|
376
|
+
|
|
377
|
+
print(summary["length"])
|
|
378
|
+
print(summary["lexical"])
|
|
379
|
+
|
|
380
|
+
print(
|
|
381
|
+
eda.vocabulary(top_n=10)
|
|
382
|
+
)
|
|
383
|
+
|
|
384
|
+
eda.plot_dataset(
|
|
385
|
+
save_path="text_dashboard.png"
|
|
386
|
+
)
|
|
387
|
+
|
|
388
|
+
eda.report(
|
|
389
|
+
"viseda_text_report.html"
|
|
390
|
+
)
|
|
391
|
+
```
|
|
392
|
+
|
|
393
|
+
TextEDA also supports directory datasets and structured files such as CSV, TSV, JSON, JSONL, Markdown, HTML, and plain text.
|
|
394
|
+
|
|
395
|
+
---
|
|
396
|
+
|
|
397
|
+
# Command-Line Interface
|
|
398
|
+
|
|
399
|
+
VisEDA installs a unified command:
|
|
400
|
+
|
|
401
|
+
```bash
|
|
402
|
+
viseda --help
|
|
403
|
+
```
|
|
404
|
+
|
|
405
|
+
Available subcommands:
|
|
406
|
+
|
|
407
|
+
```text
|
|
408
|
+
viseda image ...
|
|
409
|
+
viseda video ...
|
|
410
|
+
viseda hyper ...
|
|
411
|
+
viseda cloud ...
|
|
412
|
+
viseda text ...
|
|
413
|
+
```
|
|
414
|
+
|
|
415
|
+
### Image dataset
|
|
416
|
+
|
|
417
|
+
```bash
|
|
418
|
+
viseda image "C:/datasets/images" \
|
|
419
|
+
--label-from-parent \
|
|
420
|
+
--plot \
|
|
421
|
+
--report viseda_image_report.html
|
|
422
|
+
```
|
|
423
|
+
|
|
424
|
+
### Video dataset
|
|
425
|
+
|
|
426
|
+
```bash
|
|
427
|
+
viseda video "C:/datasets/videos" \
|
|
428
|
+
--label-from-parent \
|
|
429
|
+
--plot \
|
|
430
|
+
--report viseda_video_report.html
|
|
431
|
+
```
|
|
432
|
+
|
|
433
|
+
### Hyperspectral dataset
|
|
434
|
+
|
|
435
|
+
```bash
|
|
436
|
+
viseda hyper "C:/datasets/hyper" \
|
|
437
|
+
--label-from-parent \
|
|
438
|
+
--plot \
|
|
439
|
+
--dataset-plot \
|
|
440
|
+
--report viseda_hyper_report.html
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
### Point-cloud dataset
|
|
444
|
+
|
|
445
|
+
```bash
|
|
446
|
+
viseda cloud "C:/datasets/pointclouds" \
|
|
447
|
+
--label-from-parent \
|
|
448
|
+
--plot \
|
|
449
|
+
--report viseda_pointcloud_report.html
|
|
450
|
+
```
|
|
451
|
+
|
|
452
|
+
### Text dataset
|
|
453
|
+
|
|
454
|
+
```bash
|
|
455
|
+
viseda text "C:/datasets/text" \
|
|
456
|
+
--label-from-parent \
|
|
457
|
+
--plot \
|
|
458
|
+
--report viseda_text_report.html
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
For command-specific options:
|
|
462
|
+
|
|
463
|
+
```bash
|
|
464
|
+
viseda image --help
|
|
465
|
+
viseda video --help
|
|
466
|
+
viseda hyper --help
|
|
467
|
+
viseda cloud --help
|
|
468
|
+
viseda text --help
|
|
469
|
+
```
|
|
470
|
+
|
|
471
|
+
---
|
|
472
|
+
|
|
473
|
+
# HTML Reports
|
|
474
|
+
|
|
475
|
+
Every VisEDA module can generate a standalone HTML report.
|
|
476
|
+
|
|
477
|
+
```python
|
|
478
|
+
eda.report("report.html")
|
|
479
|
+
```
|
|
480
|
+
|
|
481
|
+
Reports are designed to provide a portable summary of dataset properties and quality indicators that can be opened directly in a browser.
|
|
482
|
+
|
|
483
|
+
Typical report sections include:
|
|
484
|
+
|
|
485
|
+
- dataset inventory
|
|
486
|
+
- labels/classes
|
|
487
|
+
- spatial or structural statistics
|
|
488
|
+
- quality metrics
|
|
489
|
+
- modality-specific statistics
|
|
490
|
+
- duplicate/similarity information where applicable
|
|
491
|
+
- summary distributions
|
|
492
|
+
- analysis metadata
|
|
493
|
+
|
|
494
|
+
---
|
|
495
|
+
|
|
496
|
+
# Visualisation Workflows
|
|
497
|
+
|
|
498
|
+
The dataset-level dashboard method is module-specific:
|
|
499
|
+
|
|
500
|
+
```python
|
|
501
|
+
# ImageEDA
|
|
502
|
+
image_eda.plot()
|
|
503
|
+
|
|
504
|
+
# VideoEDA
|
|
505
|
+
video_eda.plot_dataset()
|
|
506
|
+
|
|
507
|
+
# HyperspectralEDA
|
|
508
|
+
hyper_eda.plot_dataset()
|
|
509
|
+
|
|
510
|
+
# PointCloudEDA
|
|
511
|
+
cloud_eda.plot_dataset()
|
|
512
|
+
|
|
513
|
+
# TextEDA
|
|
514
|
+
text_eda.plot_dataset()
|
|
515
|
+
```
|
|
516
|
+
|
|
517
|
+
Specialised plotting methods are also available for modality-specific tasks such as:
|
|
518
|
+
|
|
519
|
+
- sample previews
|
|
520
|
+
- colour analysis
|
|
521
|
+
- quality analysis
|
|
522
|
+
- temporal analysis
|
|
523
|
+
- spectral analysis
|
|
524
|
+
- PCA analysis
|
|
525
|
+
- point-cloud geometry
|
|
526
|
+
- vocabulary and n-gram frequency
|
|
527
|
+
- pairwise similarity/distance matrices
|
|
528
|
+
|
|
529
|
+
---
|
|
530
|
+
|
|
531
|
+
# Supported Input Formats
|
|
532
|
+
|
|
533
|
+
| Module | Main supported inputs |
|
|
534
|
+
|---|---|
|
|
535
|
+
| **ImageEDA** | JPG/JPEG, PNG, BMP, TIFF and other supported image formats |
|
|
536
|
+
| **VideoEDA** | MP4, AVI, MOV, MKV, WEBM, MPEG/MPG, M4V |
|
|
537
|
+
| **HyperspectralEDA** | NPY, NPZ, MAT, ENVI HDR/BIL/BIP/BSQ/ENVI, TIF/TIFF |
|
|
538
|
+
| **PointCloudEDA** | NPY, NPZ, TXT, CSV, XYZ, PTS, ASCII PLY, LAS/LAZ |
|
|
539
|
+
| **TextEDA** | TXT/TEXT, MD, RST, LOG, HTML/HTM, CSV/TSV, JSON, JSONL/NDJSON |
|
|
540
|
+
|
|
541
|
+
Some formats require optional dependencies.
|
|
542
|
+
|
|
543
|
+
---
|
|
544
|
+
|
|
545
|
+
# Typical Workflow
|
|
546
|
+
|
|
547
|
+
```text
|
|
548
|
+
Raw Dataset
|
|
549
|
+
│
|
|
550
|
+
▼
|
|
551
|
+
Load with VisEDA
|
|
552
|
+
│
|
|
553
|
+
▼
|
|
554
|
+
Dataset Inventory
|
|
555
|
+
│
|
|
556
|
+
▼
|
|
557
|
+
Quality & Distribution Analysis
|
|
558
|
+
│
|
|
559
|
+
▼
|
|
560
|
+
Duplicate / Similarity Checks
|
|
561
|
+
│
|
|
562
|
+
▼
|
|
563
|
+
Visual Diagnostics
|
|
564
|
+
│
|
|
565
|
+
▼
|
|
566
|
+
HTML Report
|
|
567
|
+
│
|
|
568
|
+
▼
|
|
569
|
+
Preprocessing Decisions
|
|
570
|
+
│
|
|
571
|
+
▼
|
|
572
|
+
Model Training
|
|
573
|
+
```
|
|
574
|
+
|
|
575
|
+
---
|
|
576
|
+
|
|
577
|
+
# Package Structure
|
|
578
|
+
|
|
579
|
+
```text
|
|
580
|
+
viseda/
|
|
581
|
+
├── image/
|
|
582
|
+
├── video/
|
|
583
|
+
├── hyperspectral/
|
|
584
|
+
├── pointcloud/
|
|
585
|
+
├── text/
|
|
586
|
+
├── report/
|
|
587
|
+
├── core/
|
|
588
|
+
├── utils/
|
|
589
|
+
├── cli.py
|
|
590
|
+
└── __init__.py
|
|
591
|
+
```
|
|
592
|
+
|
|
593
|
+
The main classes can be imported directly:
|
|
594
|
+
|
|
595
|
+
```python
|
|
596
|
+
from viseda import (
|
|
597
|
+
ImageEDA,
|
|
598
|
+
VideoEDA,
|
|
599
|
+
HyperspectralEDA,
|
|
600
|
+
PointCloudEDA,
|
|
601
|
+
TextEDA,
|
|
602
|
+
)
|
|
603
|
+
```
|
|
604
|
+
|
|
605
|
+
---
|
|
606
|
+
|
|
607
|
+
# Development
|
|
608
|
+
|
|
609
|
+
Clone the repository and install in editable mode:
|
|
610
|
+
|
|
611
|
+
```bash
|
|
612
|
+
python -m pip install -e ".[dev]"
|
|
613
|
+
```
|
|
614
|
+
|
|
615
|
+
Run the tests:
|
|
616
|
+
|
|
617
|
+
```bash
|
|
618
|
+
python -m pytest
|
|
619
|
+
```
|
|
620
|
+
|
|
621
|
+
Build a distribution:
|
|
622
|
+
|
|
623
|
+
```bash
|
|
624
|
+
python -m pip install --upgrade build twine
|
|
625
|
+
|
|
626
|
+
python -m build
|
|
627
|
+
python -m twine check dist/*
|
|
628
|
+
```
|
|
629
|
+
|
|
630
|
+
---
|
|
631
|
+
|
|
632
|
+
# Documentation
|
|
633
|
+
|
|
634
|
+
The VisEDA documentation covers all five modules in detail, including:
|
|
635
|
+
|
|
636
|
+
- installation
|
|
637
|
+
- quick-start workflows
|
|
638
|
+
- complete metric explanations
|
|
639
|
+
- visualisation interpretation
|
|
640
|
+
- Python API
|
|
641
|
+
- command-line interface
|
|
642
|
+
- HTML reports
|
|
643
|
+
- test workflows
|
|
644
|
+
- advanced use cases
|
|
645
|
+
- troubleshooting
|
|
646
|
+
- API reference
|
|
647
|
+
- changelog and roadmap
|
|
648
|
+
|
|
649
|
+
---
|
|
650
|
+
|
|
651
|
+
# Citation
|
|
652
|
+
|
|
653
|
+
If you use **VisEDA** in your research, publication, dissertation, thesis, teaching, or software project, please cite the software as follows:
|
|
654
|
+
|
|
655
|
+
> **I.O.Agyemang, D. Acheampong, A.A.Baffour, & I. Adjei-Mensah**
|
|
656
|
+
> *VisEDA: A Unified Visual Exploratory Data Analysis Toolkit for Image, Video, Hyperspectral, Point Cloud, and Text Data.*
|
|
657
|
+
> GitHub repository: https://github.com/Isaac45/VisEDA
|
|
658
|
+
|
|
659
|
+
## Authors
|
|
660
|
+
|
|
661
|
+
**Isaac Osei Agyemang**
|
|
662
|
+
Data Science and Big Data Technology, Stirling College, Chengdu University, Chengdu 610054, P.R. China
|
|
663
|
+
|
|
664
|
+
**Daniel Acheampong**
|
|
665
|
+
Lutgert College, Florida Gulf Coast University, USA
|
|
666
|
+
|
|
667
|
+
**Adu Asare Baffour**
|
|
668
|
+
School of Science and Engineering, University of Missouri-Kansas City, USA
|
|
669
|
+
|
|
670
|
+
**Isaac Adjei-Mensah**
|
|
671
|
+
College of Artificial Intelligence, Yango University, Fuzhou 350015, P.R. China
|
|
672
|
+
Fujian University Engineering Research Center of Spatial Data Mining and Applications, Yango University, Fuzhou 350015, P.R. China
|
|
673
|
+
|
|
674
|
+
## BibTeX
|
|
675
|
+
|
|
676
|
+
```bibtex
|
|
677
|
+
@software{agyemang_viseda,
|
|
678
|
+
author = {
|
|
679
|
+
Agyemang, Isaac Osei and
|
|
680
|
+
Acheampong, Daniel and
|
|
681
|
+
Baffour, Adu Asare and
|
|
682
|
+
Adjei-Mensah, Isaac
|
|
683
|
+
},
|
|
684
|
+
title = {
|
|
685
|
+
VisEDA: A Unified Exploratory Data Analysis Toolkit for
|
|
686
|
+
Image, Video, Hyperspectral, Point Cloud, and Text Data
|
|
687
|
+
},
|
|
688
|
+
url = {https://github.com/Isaac45/VisEDA},
|
|
689
|
+
note = {Python software library}
|
|
690
|
+
}
|
|
691
|
+
```
|
|
692
|
+
|
|
693
|
+
> If a DOI or peer-reviewed VisEDA publication becomes available, this section can be updated with the formal publication details.
|
|
694
|
+
|
|
695
|
+
---
|
|
696
|
+
|
|
697
|
+
# Contributing
|
|
698
|
+
|
|
699
|
+
Contributions are welcome.
|
|
700
|
+
|
|
701
|
+
Useful contributions include:
|
|
702
|
+
|
|
703
|
+
- bug reports
|
|
704
|
+
- additional file-format support
|
|
705
|
+
- new EDA metrics
|
|
706
|
+
- additional visualisations
|
|
707
|
+
- documentation improvements
|
|
708
|
+
- performance optimisations
|
|
709
|
+
- tests and reproducible examples
|
|
710
|
+
|
|
711
|
+
Before proposing major API changes, please open an issue describing the motivation and expected behaviour.
|
|
712
|
+
|
|
713
|
+
---
|
|
714
|
+
|
|
715
|
+
# License
|
|
716
|
+
|
|
717
|
+
VisEDA is released under the **MIT License**.
|
|
718
|
+
|
|
719
|
+
See `LICENSE` for details.
|
|
720
|
+
|
|
721
|
+
---
|
|
722
|
+
|
|
723
|
+
<div align="center">
|
|
724
|
+
|
|
725
|
+
### VisEDA
|
|
726
|
+
|
|
727
|
+
**Explore first. Model second.**
|
|
728
|
+
|
|
729
|
+
ImageEDA · VideoEDA · HyperspectralEDA · PointCloudEDA · TextEDA
|
|
730
|
+
|
|
731
|
+
</div>
|