ethoscopy 2.1.0__tar.gz → 2.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (82) hide show
  1. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/Dockerfile +10 -1
  2. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/PKG-INFO +2 -1
  3. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/pyproject.toml +2 -1
  4. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_core.py +12 -9
  5. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_plotly.py +9 -2
  6. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_seaborn.py +2 -3
  7. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/load.py +129 -115
  8. ethoscopy-2.1.0/.claude/settings.local.json +0 -48
  9. ethoscopy-2.1.0/.coverage +0 -0
  10. ethoscopy-2.1.0/Docker/jupyterhub_data/jupyterhub.sqlite +0 -0
  11. ethoscopy-2.1.0/Docker/jupyterhub_data/jupyterhub_cookie_secret +0 -1
  12. ethoscopy-2.1.0/tasks/todo.md +0 -130
  13. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/.codecov.yml +0 -0
  14. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/.github/workflows/ci.yml +0 -0
  15. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/.github/workflows/release.yml +0 -0
  16. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/.gitignore +0 -0
  17. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/.pre-commit-config.yaml +0 -0
  18. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/CLAUDE.md +0 -0
  19. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.dummy.example +0 -0
  20. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.github.example +0 -0
  21. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.gitlab.example +0 -0
  22. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.google.example +0 -0
  23. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.keycloak +0 -0
  24. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/.env.keycloak.example +0 -0
  25. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/README.md +0 -0
  26. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/README_AUTH.md +0 -0
  27. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/config/jupyterhub_config.py +0 -0
  28. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/config/users.py +0 -0
  29. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/docker-compose.yml +0 -0
  30. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/install_r_packages.r +0 -0
  31. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/Docker/jupyterhub_data/jupyterhub_config.py +0 -0
  32. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/LICENSE +0 -0
  33. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/README.md +0 -0
  34. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/TESTING.md +0 -0
  35. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/_config.yml +0 -0
  36. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/pytest.ini +0 -0
  37. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/run_tests.py +0 -0
  38. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/scripts/README.md +0 -0
  39. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/scripts/convert_databases.sh +0 -0
  40. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/scripts/convert_wal_to_delete.py +0 -0
  41. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/scripts/publish_tutorials.py +0 -0
  42. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/setup.py +0 -0
  43. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/__init__.py +0 -0
  44. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/analyse.py +0 -0
  45. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy.py +0 -0
  46. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_HMM_class.py +0 -0
  47. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_class.py +0 -0
  48. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_draw.py +0 -0
  49. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/behavpy_periodogram_class.py +0 -0
  50. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/metadata_db.py +0 -0
  51. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/__init__.py +0 -0
  52. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/circadian_bars.py +0 -0
  53. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/general_functions.py +0 -0
  54. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/get_HMM.py +0 -0
  55. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/get_tutorials.py +0 -0
  56. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/hmm_functions.py +0 -0
  57. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/periodogram_functions.py +0 -0
  58. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/src/ethoscopy/misc/validate_datetime.py +0 -0
  59. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/__init__.py +0 -0
  60. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/conftest.py +0 -0
  61. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/data/README.md +0 -0
  62. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/data/test_ethoscope.db +0 -0
  63. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_analyse.py +0 -0
  64. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_baseline_enhancements.py +0 -0
  65. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_behavpy.py +0 -0
  66. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_behavpy_core_simple.py +0 -0
  67. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_compatibility_classes.py +0 -0
  68. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_general_functions.py +0 -0
  69. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_get_tutorials.py +0 -0
  70. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_load.py +0 -0
  71. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_load_comprehensive.py +0 -0
  72. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_load_metadata_fixes.py +0 -0
  73. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tests/test_load_optimizations.py +0 -0
  74. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/1_Overview_tutorial.ipynb +0 -0
  75. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/2_HMM_tutorial.ipynb +0 -0
  76. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/3_Circadian_tutorial.ipynb +0 -0
  77. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/4_Navigating_db_tutorial.ipynb +0 -0
  78. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/5_Ethoscopy_catch22_tutorial.ipynb +0 -0
  79. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/6_Ethoscopy_to_hctsa_tutorial.ipynb +0 -0
  80. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/ethoscope_db.csv +0 -0
  81. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/jones_et_al_metadata.csv +0 -0
  82. {ethoscopy-2.1.0 → ethoscopy-2.2.0}/tutorial_notebook/notebook_paper.ipynb +0 -0
@@ -87,9 +87,18 @@ RUN pip3 install \
87
87
  seaborn \
88
88
  bokeh \
89
89
  lifelines \
90
+ ipywidgets \
91
+ ipympl \
90
92
  opencv-python \
91
93
  mysql-connector \
92
- ethoscopy
94
+ scikit-learn \
95
+ statsmodels \
96
+ pingouin \
97
+ openpyxl \
98
+ pyarrow \
99
+ tqdm \
100
+ jupyterlab-git \
101
+ ethoscopy==2.1.0
93
102
  # pycatch22
94
103
 
95
104
  # Pre-populate tutorial datasets inside the installed package so non-root
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: ethoscopy
3
- Version: 2.1.0
3
+ Version: 2.2.0
4
4
  Summary: "A python based toolkit to download and anlyse data from the Ethoscope hardware system."
5
5
  Author-email: Lblackhurst29 <lblackhurst29@gmail.com>
6
6
  License-File: LICENSE
@@ -23,6 +23,7 @@ Requires-Dist: plotly<6.0.0,>=5.22.0
23
23
  Requires-Dist: pywavelets<2.0.0,>=1.6.0
24
24
  Requires-Dist: seaborn<0.14.0,>=0.13.2
25
25
  Requires-Dist: tabulate<0.10.0,>=0.9.0
26
+ Requires-Dist: tqdm<5.0,>=4.66
26
27
  Provides-Extra: dev
27
28
  Requires-Dist: black<25.0.0,>=24.0.0; extra == 'dev'
28
29
  Requires-Dist: ipykernel<7.0.0,>=6.29.5; extra == 'dev'
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "ethoscopy"
3
- version = "2.1.0"
3
+ version = "2.2.0"
4
4
  description = "\"A python based toolkit to download and anlyse data from the Ethoscope hardware system.\""
5
5
  authors = [{name = "Lblackhurst29",email = "lblackhurst29@gmail.com"}]
6
6
  readme = "README.md"
@@ -15,6 +15,7 @@ dependencies = [
15
15
  "pywavelets>=1.6.0,<2.0.0",
16
16
  "astropy>=7.0,<8.0",
17
17
  "nbformat>=5.10.4,<6.0.0",
18
+ "tqdm>=4.66,<5.0",
18
19
  ]
19
20
  requires-python = ">=3.12,<4.0"
20
21
  classifiers = [
@@ -9,6 +9,7 @@ from hmmlearn import hmm
9
9
  from scipy.signal import find_peaks
10
10
  from scipy.stats import zscore
11
11
  from tabulate import tabulate
12
+ from tqdm.auto import tqdm
12
13
 
13
14
  from ethoscopy.analyse import max_velocity_detector
14
15
  from ethoscopy.misc.general_functions import concat, rle
@@ -2330,9 +2331,7 @@ class behavpy_core(pd.DataFrame):
2330
2331
  seq_test = np.concatenate(test, 0)
2331
2332
  seq_test = seq_test.reshape(-1, 1)
2332
2333
 
2333
- for i in range(iterations):
2334
- print(f"Iteration {i+1} of {iterations}")
2335
-
2334
+ for i in tqdm(range(iterations), desc="HMM training", unit="iter"):
2336
2335
  init_params = ""
2337
2336
  # h = hmm.MultinomialHMM(n_components = n_states, n_iter = hmm_iterations, tol = tol, params = 'ste', verbose = verbose)
2338
2337
  h = hmm.CategoricalHMM(
@@ -2397,11 +2396,13 @@ class behavpy_core(pd.DataFrame):
2397
2396
  h.fit(seq_train, len_seq_train)
2398
2397
 
2399
2398
  # Boolean output of if the number of runs convererged on set of appropriate probabilites for s, t, an e
2400
- print(
2399
+ tqdm.write(
2401
2400
  "True Convergence:"
2402
2401
  + str(h.monitor_.history[-1] - h.monitor_.history[-2] < h.monitor_.tol)
2403
2402
  )
2404
- print("Final log liklihood score:" + str(h.score(seq_train, len_seq_train)))
2403
+ tqdm.write(
2404
+ "Final log liklihood score:" + str(h.score(seq_train, len_seq_train))
2405
+ )
2405
2406
 
2406
2407
  # On first iteration save and print the first trained matrix
2407
2408
  if i == 0:
@@ -2410,9 +2411,11 @@ class behavpy_core(pd.DataFrame):
2410
2411
  if h.score(seq_test, len_seq_test) > h_old.score(
2411
2412
  seq_test, len_seq_test
2412
2413
  ):
2413
- print("New Matrix:")
2414
+ tqdm.write("New Matrix:")
2414
2415
  df_t = pd.DataFrame(h.transmat_, index=states, columns=states)
2415
- print(tabulate(df_t, headers="keys", tablefmt="github") + "\n")
2416
+ tqdm.write(
2417
+ tabulate(df_t, headers="keys", tablefmt="github") + "\n"
2418
+ )
2416
2419
 
2417
2420
  with open(file_name, "wb") as file:
2418
2421
  pickle.dump(h, file)
@@ -2426,9 +2429,9 @@ class behavpy_core(pd.DataFrame):
2426
2429
  if h.score(seq_test, len_seq_test) > h_old.score(
2427
2430
  seq_test, len_seq_test
2428
2431
  ):
2429
- print("New Matrix:")
2432
+ tqdm.write("New Matrix:")
2430
2433
  df_t = pd.DataFrame(h.transmat_, index=states, columns=states)
2431
- print(tabulate(df_t, headers="keys", tablefmt="github") + "\n")
2434
+ tqdm.write(tabulate(df_t, headers="keys", tablefmt="github") + "\n")
2432
2435
  with open(file_name, "wb") as file:
2433
2436
  pickle.dump(h, file)
2434
2437
 
@@ -6,6 +6,7 @@ import pandas as pd
6
6
  import plotly.graph_objs as go
7
7
  from plotly.subplots import make_subplots
8
8
  from scipy.stats import zscore
9
+ from tqdm.auto import tqdm
9
10
 
10
11
  from ethoscopy.behavpy_draw import behavpy_draw
11
12
  from ethoscopy.behavpy_seaborn import behavpy_seaborn
@@ -3559,10 +3560,16 @@ class behavpy_plotly(behavpy_draw):
3559
3560
  min_t = []
3560
3561
  max_t = []
3561
3562
 
3562
- for c, (df, b) in enumerate(zip(decoded, b_list)):
3563
+ for c, (df, b) in enumerate(
3564
+ tqdm(
3565
+ list(zip(decoded, b_list)),
3566
+ desc="Plotting",
3567
+ unit="fly",
3568
+ total=len(decoded),
3569
+ )
3570
+ ):
3563
3571
  df["colour"] = df["previous_state"].map(colours_index)
3564
3572
  id = df.first_valid_index()
3565
- print(f"Plotting: {id}")
3566
3573
 
3567
3574
  if stim_df is not None:
3568
3575
  df2 = stim_df.xmv("id", id)
@@ -6,6 +6,7 @@ import matplotlib.pyplot as plt
6
6
  import numpy as np
7
7
  import pandas as pd
8
8
  import seaborn as sns
9
+ from tqdm.auto import tqdm
9
10
 
10
11
  from ethoscopy.behavpy_draw import behavpy_draw
11
12
  from ethoscopy.misc.circadian_bars import circadian_bars
@@ -3000,9 +3001,7 @@ class behavpy_seaborn(behavpy_draw):
3000
3001
  min_t = []
3001
3002
  max_t = []
3002
3003
 
3003
- for ax, data in enumerate(decoded):
3004
- print(f"Plotting: {data.index[0]}")
3005
-
3004
+ for ax, data in enumerate(tqdm(decoded, desc="Plotting", unit="fly")):
3006
3005
  data["bin"] = data["bin"] / (60 * 60)
3007
3006
  st = data["state"].to_numpy()
3008
3007
  time = data["bin"].to_numpy()
@@ -15,6 +15,7 @@ from urllib.parse import urlparse
15
15
 
16
16
  import numpy as np
17
17
  import pandas as pd
18
+ from tqdm.auto import tqdm
18
19
 
19
20
  from ethoscopy.misc.validate_datetime import validate_datetime
20
21
 
@@ -74,7 +75,7 @@ def _connect_db(path):
74
75
  return sqlite3.connect(path_str)
75
76
 
76
77
 
77
- def download_from_remote_dir(meta, remote_dir, local_dir):
78
+ def download_from_remote_dir(meta, remote_dir, local_dir, progress=True):
78
79
  """
79
80
  Download ethoscope data from a remote FTP server to a local directory.
80
81
 
@@ -88,6 +89,8 @@ def download_from_remote_dir(meta, remote_dir, local_dir):
88
89
  e.g. 'ftp://YOUR_SERVER//auto_generated_data//ethoscope_results'
89
90
  local_dir (str): Path to the local directory for saving .db files. Files will be saved using the FTP server's structure.
90
91
  e.g. 'C:\\Users\\YOUR_NAME\\Documents\\ethoscope_databases'
92
+ progress (bool, optional): If True, show a tqdm progress bar (ipywidgets-based in Jupyter,
93
+ text in CLI). Default is True.
91
94
 
92
95
  Returns:
93
96
  None
@@ -268,39 +271,27 @@ def download_from_remote_dir(meta, remote_dir, local_dir):
268
271
  ftp.quit()
269
272
  localfile.close()
270
273
 
271
- # iterate over paths, downloading each file
272
- # provide estimate download time based upon average time of previous downloads in queue
274
+ # iterate over paths, downloading each file. tqdm provides per-file progress
275
+ # plus an aggregate ETA, so we no longer need to hand-roll an estimate.
273
276
  download = partial(
274
277
  download_database,
275
278
  remote_dir=parse.netloc,
276
279
  folders=parse.path,
277
280
  local_dir=local_dir,
278
281
  )
279
- times = []
280
282
 
281
- for counter, j in enumerate(paths):
282
- print(
283
- "Downloading {}... {}/{}".format(
284
- j[0].split("/")[1], counter + 1, len(paths)
285
- )
286
- )
287
- if counter == 0:
288
- start = time.time()
289
- p = PurePosixPath(j[0])
290
- download(work_dir=p.parents[0], file_name=p.name, file_size=j[1])
291
- stop = time.time()
292
- t = stop - start
293
- times.append(t)
294
-
295
- else:
296
- av_time = round((np.mean(times) / 60) * (len(paths) - (counter + 1)))
297
- print(f"Estimated finish time: {av_time} mins")
298
- start = time.time()
299
- p = PurePosixPath(j[0])
300
- download(work_dir=p.parents[0], file_name=p.name, file_size=j[1])
301
- stop = time.time()
302
- t = stop - start
303
- times.append(t)
283
+ iterator = tqdm(
284
+ paths,
285
+ desc="Downloading databases",
286
+ unit="db",
287
+ disable=not progress,
288
+ )
289
+ for j in iterator:
290
+ machine_name = j[0].split("/")[1]
291
+ if progress:
292
+ iterator.set_postfix_str(machine_name, refresh=False)
293
+ p = PurePosixPath(j[0])
294
+ download(work_dir=p.parents[0], file_name=p.name, file_size=j[1])
304
295
 
305
296
 
306
297
  def link_meta_index(metadata, local_dir):
@@ -471,6 +462,7 @@ def load_ethoscope(
471
462
  cache=None,
472
463
  FUN=None,
473
464
  verbose=True,
465
+ progress=True,
474
466
  ):
475
467
  """
476
468
  Load and process ethoscope data from database files.
@@ -488,7 +480,9 @@ def load_ethoscope(
488
480
  Directory structure mirrors ethoscope saved data. Cached files are in pickle format. Default is None.
489
481
  FUN (callable, optional): Function to apply individual curation to each ROI, typically using package
490
482
  generated functions (e.g., sleep_annotation). If None, data remains as found in the database. Default is None.
491
- verbose (bool, optional): If True, prints information about each ROI when loading. Default is True.
483
+ verbose (bool, optional): If True, emits per-ROI warnings when loading fails. Default is True.
484
+ progress (bool, optional): If True, show a tqdm progress bar over ROIs (ipywidgets-based in
485
+ Jupyter, text in CLI). Default is True.
492
486
 
493
487
  Returns:
494
488
  pd.DataFrame: DataFrame containing the database data with unique IDs per fly as the index
@@ -507,101 +501,114 @@ def load_ethoscope(
507
501
  # Group ROIs by database file to reuse connections and cache metadata
508
502
  grouped_metadata = metadata.groupby("path")
509
503
 
504
+ pbar = tqdm(
505
+ total=len(metadata),
506
+ desc="Loading ROIs",
507
+ unit="roi",
508
+ disable=not progress,
509
+ )
510
+
510
511
  # iterate over each database file
511
- for db_path, group in grouped_metadata:
512
- conn = None
512
+ try:
513
+ for db_path, group in grouped_metadata:
514
+ conn = None
513
515
 
514
- try:
515
- # Open connection once per database file
516
- conn = _connect_db(db_path)
517
-
518
- # Cache metadata queries that are the same for all ROIs in this database
519
- roi_df = pd.read_sql_query("SELECT * FROM ROI_MAP", conn)
520
- var_df = pd.read_sql_query("SELECT * FROM VAR_MAP", conn)
521
- date = pd.read_sql_query(
522
- 'SELECT value FROM METADATA WHERE field = "date_time"', conn
523
- )
524
- if date.empty:
525
- raise ValueError("No date_time found in METADATA table")
526
- date_formatted = time.strftime(
527
- "%Y-%m-%d %H:%M:%S", time.gmtime(float(date.iloc[0].iloc[0]))
528
- )
516
+ try:
517
+ # Open connection once per database file
518
+ conn = _connect_db(db_path)
519
+
520
+ # Cache metadata queries that are the same for all ROIs in this database
521
+ roi_df = pd.read_sql_query("SELECT * FROM ROI_MAP", conn)
522
+ var_df = pd.read_sql_query("SELECT * FROM VAR_MAP", conn)
523
+ date = pd.read_sql_query(
524
+ 'SELECT value FROM METADATA WHERE field = "date_time"', conn
525
+ )
526
+ if date.empty:
527
+ raise ValueError("No date_time found in METADATA table")
528
+ date_formatted = time.strftime(
529
+ "%Y-%m-%d %H:%M:%S", time.gmtime(float(date.iloc[0].iloc[0]))
530
+ )
529
531
 
530
- # Process each ROI in this database
531
- for i in group.index:
532
- file_info = metadata.iloc[metadata.index.get_loc(i), :]
532
+ # Process each ROI in this database
533
+ for i in group.index:
534
+ file_info = metadata.iloc[metadata.index.get_loc(i), :]
533
535
 
534
- try:
535
- if verbose is True:
536
- print(
537
- "Loading ROI_{} from {}".format(
538
- file_info["region_id"], file_info["machine_name"]
536
+ try:
537
+ if progress:
538
+ pbar.set_postfix_str(
539
+ f"{file_info['machine_name']} ROI_{file_info['region_id']}",
540
+ refresh=False,
539
541
  )
540
- )
541
542
 
542
- # Use optimized single ROI reader with cached connection and metadata
543
- roi_1 = read_single_roi_optimized(
544
- file_info,
545
- conn,
546
- roi_df,
547
- var_df,
548
- date_formatted,
549
- min_time,
550
- max_time,
551
- reference_hour,
552
- cache,
553
- )
543
+ # Use optimized single ROI reader with cached connection and metadata
544
+ roi_1 = read_single_roi_optimized(
545
+ file_info,
546
+ conn,
547
+ roi_df,
548
+ var_df,
549
+ date_formatted,
550
+ min_time,
551
+ max_time,
552
+ reference_hour,
553
+ cache,
554
+ )
554
555
 
555
- if roi_1 is None:
556
- if verbose is True:
557
- print(
558
- "ROI_{} from {} was unable to load due to an error formatting roi".format(
559
- file_info["region_id"], file_info["machine_name"]
556
+ if roi_1 is None:
557
+ if verbose is True:
558
+ tqdm.write(
559
+ "ROI_{} from {} was unable to load due to an error formatting roi".format(
560
+ file_info["region_id"],
561
+ file_info["machine_name"],
562
+ )
560
563
  )
561
- )
562
- continue
564
+ continue
565
+
566
+ if FUN is not None:
567
+ roi_1 = FUN(roi_1)
568
+
569
+ if roi_1 is None:
570
+ if verbose is True:
571
+ tqdm.write(
572
+ "ROI_{} from {} was unable to load due to an error in applying the function".format(
573
+ file_info["region_id"],
574
+ file_info["machine_name"],
575
+ )
576
+ )
577
+ continue
563
578
 
564
- if FUN is not None:
565
- roi_1 = FUN(roi_1)
579
+ # Check if 'id' column already exists, if not insert it
580
+ if "id" not in roi_1.columns:
581
+ roi_1.insert(0, "id", file_info["id"])
582
+ else:
583
+ # Replace existing id with the one from metadata for consistency
584
+ roi_1["id"] = file_info["id"]
566
585
 
567
- if roi_1 is None:
586
+ # Add to list instead of concatenating in loop
587
+ roi_data_list.append(roi_1)
588
+
589
+ except Exception as e:
568
590
  if verbose is True:
569
- print(
570
- "ROI_{} from {} was unable to load due to an error in applying the function".format(
571
- file_info["region_id"], file_info["machine_name"]
591
+ tqdm.write(
592
+ "ROI_{} from {} was unable to load due to an error loading roi: {}".format(
593
+ file_info["region_id"],
594
+ file_info["machine_name"],
595
+ str(e),
572
596
  )
573
597
  )
574
- continue
575
-
576
- # Check if 'id' column already exists, if not insert it
577
- if "id" not in roi_1.columns:
578
- roi_1.insert(0, "id", file_info["id"])
579
- else:
580
- # Replace existing id with the one from metadata for consistency
581
- roi_1["id"] = file_info["id"]
582
-
583
- # Add to list instead of concatenating in loop
584
- roi_data_list.append(roi_1)
585
-
586
- except Exception as e:
587
- if verbose is True:
588
- print(
589
- "ROI_{} from {} was unable to load due to an error loading roi: {}".format(
590
- file_info["region_id"],
591
- file_info["machine_name"],
592
- str(e),
593
- )
594
- )
595
- import traceback
598
+ import traceback
596
599
 
597
- print("Full traceback:")
598
- traceback.print_exc()
599
- continue
600
+ tqdm.write("Full traceback:")
601
+ tqdm.write(traceback.format_exc())
602
+ continue
603
+ finally:
604
+ pbar.update(1)
600
605
 
601
- finally:
602
- # Close connection when done with this database
603
- if conn:
604
- conn.close()
606
+ finally:
607
+ # Close connection when done with this database
608
+ if conn:
609
+ conn.close()
610
+ finally:
611
+ pbar.close()
605
612
 
606
613
  # Concatenate all data at once for much better performance
607
614
  if roi_data_list:
@@ -612,7 +619,7 @@ def load_ethoscope(
612
619
  return data
613
620
 
614
621
 
615
- def load_ethoscope_metadata(metadata):
622
+ def load_ethoscope_metadata(metadata, progress=True):
616
623
  """
617
624
  Extract metadata from ethoscope database files.
618
625
 
@@ -621,6 +628,8 @@ def load_ethoscope_metadata(metadata):
621
628
 
622
629
  Args:
623
630
  metadata (pd.DataFrame): Metadata dataframe as returned from link_meta_index function
631
+ progress (bool, optional): If True, show a tqdm progress bar (ipywidgets-based in Jupyter,
632
+ text in CLI). Default is True.
624
633
 
625
634
  Returns:
626
635
  pd.DataFrame: DataFrame containing the metadata from the METADATA table in each ethoscope database,
@@ -724,7 +733,12 @@ def load_ethoscope_metadata(metadata):
724
733
  rows = []
725
734
 
726
735
  # iterate over each ethoscope in the metadata df
727
- for i in meta_df["path"]:
736
+ for i in tqdm(
737
+ meta_df["path"],
738
+ desc="Reading metadata",
739
+ unit="db",
740
+ disable=not progress,
741
+ ):
728
742
  row = get_meta(i)
729
743
  rows.append(row)
730
744
 
@@ -966,14 +980,14 @@ def read_single_roi_optimized(
966
980
  # Handle "database disk image is malformed" errors
967
981
  # This can occur with WAL-mode databases on read-only mounts
968
982
  if "malformed" in str(e).lower() or "disk image" in str(e).lower():
969
- print(
983
+ tqdm.write(
970
984
  f"Warning: Database error for ROI {file['region_id']}, attempting retry with fresh connection..."
971
985
  )
972
986
 
973
987
  # Get database path from file metadata
974
988
  db_path = file.get("path")
975
989
  if not db_path:
976
- print(
990
+ tqdm.write(
977
991
  "Error: Cannot retry - database path not found in file metadata"
978
992
  )
979
993
  raise
@@ -983,9 +997,9 @@ def read_single_roi_optimized(
983
997
  try:
984
998
  retry_conn = _connect_db(db_path)
985
999
  data = pd.read_sql_query(sql_query, retry_conn)
986
- print(f"Success: ROI {file['region_id']} loaded on retry")
1000
+ tqdm.write(f"Success: ROI {file['region_id']} loaded on retry")
987
1001
  except Exception as retry_error:
988
- print(
1002
+ tqdm.write(
989
1003
  f"Error: Retry failed for ROI {file['region_id']}: {retry_error}"
990
1004
  )
991
1005
  raise
@@ -1032,5 +1046,5 @@ def read_single_roi_optimized(
1032
1046
  return data
1033
1047
 
1034
1048
  except Exception as e:
1035
- print(f"Error reading ROI {file['region_id']}: {e}")
1049
+ tqdm.write(f"Error reading ROI {file['region_id']}: {e}")
1036
1050
  return None
@@ -1,48 +0,0 @@
1
- {
2
- "permissions": {
3
- "allow": [
4
- "Bash(wc:*)",
5
- "Bash(paplay:*)",
6
- "Bash(ETHOSCOPE_DATA_PATH=/mnt/ethoscope_data docker compose:*)",
7
- "Bash(docker ps:*)",
8
- "Bash(docker logs:*)",
9
- "Bash(docker compose:*)",
10
- "Bash(curl:*)",
11
- "Bash(docker exec:*)",
12
- "Bash(mkdir -p /tmp/ethoscopy_wheel_inspect)",
13
- "Read(//tmp/**)",
14
- "Bash(pip download *)",
15
- "Bash(unzip -l ethoscopy-2.0.4-py3-none-any.whl)",
16
- "Bash(git check-ignore *)",
17
- "Bash(/home/gg/Code/ethoscope_project/ethoscopy/.venv/bin/pip install *)",
18
- "Bash(.venv/bin/python -m pytest tests/test_get_tutorials.py -v)",
19
- "Bash(.venv/bin/python *)",
20
- "Bash(.venv/bin/pip install *)",
21
- "Bash(gh auth *)",
22
- "Bash(gh repo *)",
23
- "Bash(docker info *)",
24
- "mcp__bookstack__bookstack_search",
25
- "Bash(.venv/bin/ruff check *)",
26
- "Bash(.venv/bin/black --check src/ tests/)",
27
- "Bash(.venv/bin/black src/ethoscopy/misc/get_tutorials.py tests/test_get_tutorials.py)",
28
- "mcp__bookstack__bookstack_pages_read",
29
- "mcp__bookstack__bookstack_pages_update",
30
- "Bash(gh issue create --repo gilestrolab/ethoscopy --title 'Tutorial data pickles missing from PyPI wheel \\(2.0.0 – 2.0.4\\)' --body ' *)",
31
- "Bash(git add *)",
32
- "Bash(git commit *)",
33
- "Bash(git push *)",
34
- "Bash(git tag *)",
35
- "Bash(gh release create v2.0.5 --repo gilestrolab/ethoscopy --title 'v2.0.5 — package tutorial data separately' --notes ' *)",
36
- "Bash(gh run *)",
37
- "Bash(gh issue *)",
38
- "Bash(.venv/bin/twine check *)",
39
- "Bash(.venv/bin/twine upload *)",
40
- "Bash(.venv/bin/pip index *)",
41
- "Bash(ETHOSCOPE_LAB_TAG=1.2 docker compose build)",
42
- "Bash(gh release *)",
43
- "Bash(ETHOSCOPE_LAB_TAG=1.2 docker compose build --no-cache --progress=plain)",
44
- "Bash(grep -E --line-buffered '^#[0-9]+ \\(ERROR|CANCELED\\)|^#[0-9]+ DONE [0-9.]+s$|^ERROR|^failed to solve|Execution halted|^Error: Required R packages|non-zero code|^naming to|^ *=> .* done *$')",
45
- "Bash(grep -E --line-buffered '^#[0-9]+ \\(ERROR|CANCELED\\)$|^#[0-9]+ DONE [0-9.]+s$|^failed to solve|Execution halted|^Error: Required R packages|Successfully built|naming to docker\\\\.io')"
46
- ]
47
- }
48
- }
ethoscopy-2.1.0/.coverage DELETED
Binary file
@@ -1 +0,0 @@
1
- 28d1b26a4bf0427c5ac0a6b896c90ffedd826af2f5fce289877ae09bf4a26b75
@@ -1,130 +0,0 @@
1
- # Ethoscopy — task log
2
-
3
- ## 2026-04-21 — Fix missing tutorial pickles on PyPI installs
4
-
5
- **Problem.** Users installing ethoscopy 2.0.4 from PyPI hit
6
- `FileNotFoundError: Tutorial data files not found in: …/ethoscopy/misc/tutorial_data`
7
- when running `get_tutorial('overview')`. Root cause: `.gitignore` line 9
8
- contains `*.pkl`; Hatchling's default VCS plugin honours `.gitignore` when
9
- selecting wheel contents, so every tracked pickle was silently dropped from
10
- the 2.0.4 wheel (confirmed by inspecting the 144 KB artifact on PyPI).
11
-
12
- **Decision.** Keep the pickles *out* of the wheel by design
13
- (`overview_data.pkl` alone is ~31 MB — ~200× the code payload). Instead,
14
- ship an explicit fetch path and document it clearly everywhere.
15
-
16
- ### Changes
17
-
18
- - [x] `pyproject.toml`: add explicit `[tool.hatch.build.targets.wheel] exclude`
19
- for `src/ethoscopy/misc/tutorial_data/*.pkl` so intent is independent of
20
- `.gitignore`.
21
- - [x] `src/ethoscopy/misc/get_tutorials.py`: add `download_tutorial_data()`
22
- (stdlib `urllib`, idempotent, overwrite flag) and rewrite the
23
- `FileNotFoundError` with a copy-paste recovery snippet + GitHub URL.
24
- - [x] `src/ethoscopy/misc/get_HMM.py`: update its `FileNotFoundError` to point
25
- at the same helper (the 4-state HMM pickles live in the same folder).
26
- - [x] `src/ethoscopy/__init__.py`: re-export `download_tutorial_data` and
27
- `get_tutorial` at the top level so users call
28
- `etho.download_tutorial_data()`.
29
- - [x] `tests/test_get_tutorials.py`: add 5 tests for the new helper (URL
30
- coverage, full download, skip-if-present, overwrite, network error).
31
- 15 tests pass.
32
- - [x] Six tutorial notebooks (`1_Overview`, `2_HMM`, `3_Circadian`,
33
- `5_Ethoscopy_catch22`, `6_Ethoscopy_to_hctsa`, `notebook_paper`): insert
34
- one markdown + one code cell before each first `get_tutorial(...)`
35
- explaining the one-time fetch.
36
- - [x] `README.md`: add a "Tutorial data" section with the one-liner plus a
37
- fallback manual-download path.
38
-
39
- ### Follow-up (same session)
40
-
41
- After shipping the first cut, a second concern surfaced: the default
42
- destination (`<site-packages>/ethoscopy/misc/tutorial_data/`) is only
43
- writable for the user who installed ethoscopy. In system-wide installs,
44
- conda base envs, or Docker images where the package was installed by
45
- root, a non-root user calling `etho.download_tutorial_data()` would hit
46
- `PermissionError`. Fixed by:
47
-
48
- - [x] Default `download_tutorial_data(dest_dir=...)` to
49
- `~/.cache/ethoscopy/tutorial_data/` (always user-writable).
50
- - [x] `get_tutorial` / `get_HMM` now consult three locations in order:
51
- (1) package dir, (2) `$ETHOSCOPY_TUTORIAL_DATA_DIR`, (3) user cache.
52
- `_missing_files_message()` prints the full search list so users
53
- know exactly where ethoscopy looked.
54
- - [x] `PermissionError` on `mkdir` surfaces a wrapped message pointing
55
- at `dest_dir=` / the env override.
56
- - [x] New tests: search-path ordering, env-var override, cache fallback,
57
- permission-error wrapping, default-dest-is-user-cache. 22/22 pass.
58
- - [x] `Docker/Dockerfile` now runs
59
- `download_tutorial_data(dest_dir=package_tutorial_data_dir())`
60
- during build so JupyterHub users inherit the pickles from the
61
- root-owned package dir — no runtime download ever needed.
62
- - [x] `README.md` "Tutorial data" section documents the lookup order.
63
-
64
- ### Verification
65
-
66
- - [x] `python -m build --wheel` → 146 KB wheel, 24 files, 0 `.pkl`.
67
- - [x] Live network check of `etho.download_tutorial_data()` against
68
- GitHub raw URLs — will be exercised by the Docker build.
69
- - [x] Build the Docker image with
70
- `docker compose build` — in-flight.
71
-
72
- ### Ship checklist (session of 2026-04-21)
73
-
74
- - [x] `pyproject.toml` bumped 2.0.4 → 2.0.5.
75
- - [x] Ruff + black both clean on `src/` and `tests/`.
76
- - [x] Bookstack "Getting started" page (book 1 / page 1) updated with
77
- new "Tutorial data" section; dependency list refreshed; editor
78
- switched from `wysiwyg` → `markdown` as a side effect of using
79
- the markdown field.
80
- - [x] GitHub issue #7 opened and auto-closed by the commit footer.
81
- - [x] Commit `c858b84` pushed to `main` (13 files, +593 / −37).
82
- - [x] Tag `v2.0.5` pushed.
83
- - [x] GitHub release v2.0.5 published.
84
- - [x] CI `release.yml` failed on both tag-push and release events —
85
- **pre-existing** `tests/test_load_optimizations.py` failures
86
- (`No date_time found in METADATA table`), same pattern as 2.0.4.
87
- Unrelated to this change.
88
- - [x] Manual `twine upload` to PyPI succeeded: 2.0.5 wheel (147 KB)
89
- and sdist (14 MB) live at
90
- <https://pypi.org/project/ethoscopy/2.0.5/>.
91
- - [ ] Docker image `ggilestro/ethoscope-lab:1.2` build + push.
92
-
93
- ### Side debt spotted during this session
94
-
95
- - `tests/test_load_optimizations.py::TestLoadOptimizationPerformance`
96
- (`test_load_ethoscope_memory_usage`, `test_connection_caching_benefit`)
97
- fails on GitHub Actions for both 2.0.4 and 2.0.5 tag / release runs.
98
- Raises `ValueError: No date_time found in METADATA table` from
99
- `load.py:524`. The test fixture likely builds a SQLite DB without the
100
- METADATA row that `load_ethoscope` now requires after
101
- a6a1473 (`harden load_ethoscope_metadata against firmware quirks`).
102
- Fix: update the test fixture to seed a `date_time` value in METADATA.
103
- - `.github/workflows/release.yml` also runs broader-than-unit tests on
104
- tag push (`pytest tests/ -v --cov=ethoscopy -m "not slow"`), which is
105
- stricter than `ci.yml`. This is why the `test` job blocks automated
106
- PyPI publish. Either tighten the selection to `-m unit` or fix the
107
- underlying test.
108
- - `.gitignore`'s blanket `*.pkl` is still present. It's currently
109
- harmless because `pyproject.toml` has an explicit wheel exclusion,
110
- but a future maintainer unaware of the wheel-build behaviour could
111
- be surprised.
112
-
113
- ### Discovered during work
114
-
115
- - `.gitignore` has a blanket `*.pkl` rule that's also generating noise (it
116
- affects the behaviour of `git add` for tutorial pickles and the Hatchling
117
- wheel). The explicit `exclude` in `pyproject.toml` now removes any ambiguity
118
- about packaging, but the `.gitignore` rule could be tightened to a
119
- user-output pattern (e.g. `tutorial_dataframe.pkl`) in a future pass.
120
- - The project-local `.venv` had a stale shebang (pointing to
121
- `/home/gg/Data/...`) and had to be rebuilt.
122
-
123
- ### Review
124
-
125
- - Net diff intentionally small: one new helper (~40 lines), one packaging
126
- guard, one README section, one notebook preamble (templated, six files).
127
- - No behaviour change for users who already have the pickles on disk; the
128
- error path now self-documents recovery.
129
- - PyPI release 2.0.5 should carry these changes — the 2.0.4 wheel remains
130
- broken for new installers until then.
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes