diffindiff 2.5.5__tar.gz → 2.5.6__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (21) hide show
  1. {diffindiff-2.5.5 → diffindiff-2.5.6}/PKG-INFO +7 -6
  2. {diffindiff-2.5.5 → diffindiff-2.5.6}/README.md +5 -5
  3. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/config.py +3 -3
  4. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/didtools.py +118 -77
  5. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff.egg-info/PKG-INFO +7 -6
  6. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff.egg-info/requires.txt +4 -2
  7. {diffindiff-2.5.5 → diffindiff-2.5.6}/setup.py +7 -3
  8. {diffindiff-2.5.5 → diffindiff-2.5.6}/MANIFEST.in +0 -0
  9. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/__init__.py +0 -0
  10. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/didanalysis.py +0 -0
  11. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/didanalysis_helper.py +0 -0
  12. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/diddata.py +0 -0
  13. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/tests/__init__.py +0 -0
  14. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/tests/data/Corona_Hesse.xlsx +0 -0
  15. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/tests/data/counties_DE.csv +0 -0
  16. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/tests/data/curfew_DE.csv +0 -0
  17. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff/tests/tests_diffindiff.py +0 -0
  18. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff.egg-info/SOURCES.txt +0 -0
  19. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff.egg-info/dependency_links.txt +0 -0
  20. {diffindiff-2.5.5 → diffindiff-2.5.6}/diffindiff.egg-info/top_level.txt +0 -0
  21. {diffindiff-2.5.5 → diffindiff-2.5.6}/setup.cfg +0 -0
@@ -1,10 +1,11 @@
1
1
  Metadata-Version: 2.1
2
2
  Name: diffindiff
3
- Version: 2.5.5
3
+ Version: 2.5.6
4
4
  Summary: diffindiff: Python library for convenient Difference-in-Differences analyses
5
5
  Author: Thomas Wieland
6
6
  Author-email: geowieland@googlemail.com
7
7
  Description-Content-Type: text/markdown
8
+ Provides-Extra: optional
8
9
 
9
10
  # diffindiff: Python library for convenient Difference-in-Differences analyses
10
11
 
@@ -29,7 +30,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
29
30
 
30
31
  If you use this software, please cite:
31
32
 
32
- Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.5) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
33
+ Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.6) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
33
34
 
34
35
 
35
36
  ## Installation
@@ -177,11 +178,11 @@ See the /tests directory for usage examples of most of the included functions.
177
178
 
178
179
  ## AI Usage Statement
179
180
 
180
- This software was developed without the use of AI-generated code. The Continue Agent in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
181
+ This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
181
182
 
182
183
 
183
- ## What's new (v2.5.5)
184
+ ## What's new (v2.5.6)
184
185
 
185
186
  - Bugfixes
186
- - didanalysis.did_analysis(): Correct usage of param 'log_outcome' in any case and check whether value of 'log_outcome_add' is appropriate
187
- - didanalysis.DiffModel.plot() and didanalysis.DiffModel.plot_counterfactual(): Correct detection of the date format via object metadata
187
+ - didtools.is_parallel(): Checking whether there is a pre-treatment period and skipping test if not (additional NOTE)
188
+ - Optional installation of XGBoost and LightGBM (more stable if problems with the installation of these packages occur)
@@ -21,7 +21,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
21
21
 
22
22
  If you use this software, please cite:
23
23
 
24
- Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.5) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
24
+ Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.6) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
25
25
 
26
26
 
27
27
  ## Installation
@@ -169,11 +169,11 @@ See the /tests directory for usage examples of most of the included functions.
169
169
 
170
170
  ## AI Usage Statement
171
171
 
172
- This software was developed without the use of AI-generated code. The Continue Agent in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
172
+ This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
173
173
 
174
174
 
175
- ## What's new (v2.5.5)
175
+ ## What's new (v2.5.6)
176
176
 
177
177
  - Bugfixes
178
- - didanalysis.did_analysis(): Correct usage of param 'log_outcome' in any case and check whether value of 'log_outcome_add' is appropriate
179
- - didanalysis.DiffModel.plot() and didanalysis.DiffModel.plot_counterfactual(): Correct detection of the date format via object metadata
178
+ - didtools.is_parallel(): Checking whether there is a pre-treatment period and skipping test if not (additional NOTE)
179
+ - Optional installation of XGBoost and LightGBM (more stable if problems with the installation of these packages occur)
@@ -4,15 +4,15 @@
4
4
  # Author: Thomas Wieland
5
5
  # ORCID: 0000-0001-5168-9846
6
6
  # mail: geowieland@googlemail.com
7
- # Version: 1.0.26
8
- # Last update: 2026-09-11 21:46
7
+ # Version: 1.0.27
8
+ # Last update: 2026-10-03 11:03
9
9
  # Copyright (c) 2025-2026 Thomas Wieland
10
10
  #-----------------------------------------------------------------------
11
11
 
12
12
  # Basic config:
13
13
 
14
14
  PACKAGE_NAME = "diffindiff"
15
- PACKAGE_VERSION = "2.5.5"
15
+ PACKAGE_VERSION = "2.5.6"
16
16
 
17
17
  VERBOSE = False
18
18
 
@@ -4,8 +4,8 @@
4
4
  # Author: Thomas Wieland
5
5
  # ORCID: 0000-0001-5168-9846
6
6
  # mail: geowieland@googlemail.com
7
- # Version: 2.2.5
8
- # Last update: 2026-07-23 19:34
7
+ # Version: 2.2.6
8
+ # Last update: 2026-10-03 11:02
9
9
  # Copyright (c) 2025-2026 Thomas Wieland
10
10
  #-----------------------------------------------------------------------
11
11
 
@@ -22,11 +22,20 @@ from sklearn.svm import SVR
22
22
  from sklearn.neighbors import KNeighborsRegressor
23
23
  from sklearn.pipeline import Pipeline
24
24
  from sklearn.preprocessing import StandardScaler
25
- from xgboost import XGBRegressor
26
- from lightgbm import LGBMRegressor
27
25
  from sklearn.linear_model import LinearRegression
28
26
  from sklearn.model_selection import train_test_split
29
27
  from sklearn.neural_network import MLPRegressor
28
+
29
+ try:
30
+ from xgboost import XGBRegressor
31
+ except ImportError:
32
+ XGBRegressor = None
33
+ try:
34
+ from lightgbm import LGBMRegressor
35
+ except ImportError:
36
+ LGBMRegressor = None
37
+
38
+
30
39
  import diffindiff.config as config
31
40
 
32
41
 
@@ -826,83 +835,105 @@ def is_parallel(
826
835
  treatment_col = treatment_col,
827
836
  verbose = False
828
837
  )
829
-
830
- if verbose:
831
- print(f"Testing outcome '{outcome_col}' for parallel time trends", end = " ... ")
832
-
833
- if pre_post or not modeldata_isnotreatment:
834
- parallel = "not_tested"
835
- test_ols_model = None
836
-
837
- treatment_group = modeldata_isnotreatment[1]
838
838
 
839
- if config.ACCEPT_CONTINUOUS_TREATMENTS:
840
-
841
- if len(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] > 0)]) > 0:
842
-
843
- first_day_of_treatment = min(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] > 0)][time_col])
844
-
845
- data_test = data[data[time_col] < first_day_of_treatment].copy()
846
- data_test[config.TG_COL] = 0
847
- data_test.loc[data_test[unit_col].isin(treatment_group), config.TG_COL] = 1
848
-
849
- if config.TIME_COUNTER_COL not in data_test.columns:
850
- data_test = date_counter(
851
- df = data_test,
852
- date_col = time_col,
853
- new_col = config.TIME_COUNTER_COL,
854
- verbose = False
855
- )
856
- data_test[f"{config.TG_COL}_x_{config.TIME_COL}"] = data_test[config.TG_COL]*data_test[config.TIME_COUNTER_COL]
839
+ parallel = "not_tested"
840
+ test_ols_model = None
857
841
 
858
- test_ols_model = ols(f'{outcome_col} ~ {config.TG_COL} + {config.TIME_COUNTER_COL} + {config.TG_COL}_x_{config.TIME_COL}', data = data_test).fit()
859
- coef_TG_x_t_p = test_ols_model.pvalues[f"{config.TG_COL}_x_{config.TIME_COL}"]
842
+ if not pre_post:
860
843
 
861
- if coef_TG_x_t_p < alpha:
862
- parallel = False
863
- else:
864
- parallel = True
844
+ no_pre_period = False
865
845
 
866
- else:
867
- parallel = "not_tested"
868
- test_ols_model = None
846
+ if verbose:
847
+ print(f"Testing outcome '{outcome_col}' for parallel time trends", end = " ... ")
869
848
 
870
- else:
849
+ treatment_group = modeldata_isnotreatment[1]
871
850
 
872
- if len(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] == 1)]) > 0:
873
-
874
- first_day_of_treatment = min(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] == 1)][time_col])
851
+ if config.ACCEPT_CONTINUOUS_TREATMENTS:
875
852
 
876
- data_test = data[data[time_col] < first_day_of_treatment].copy()
877
- data_test[config.TG_COL] = 0
878
- data_test.loc[data_test[unit_col].isin(treatment_group), config.TG_COL] = 1
853
+ if len(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] > 0)]) > 0:
879
854
 
880
- if config.TIME_COUNTER_COL not in data_test.columns:
881
- data_test = date_counter(
882
- df = data_test,
883
- date_col = time_col,
884
- new_col = config.TIME_COUNTER_COL,
885
- verbose = False
886
- )
887
- data_test[f"{config.TG_COL}_x_{config.TIME_COL}"] = data_test[config.TG_COL]*data_test[config.TIME_COUNTER_COL]
855
+ first_day_of_treatment = min(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] > 0)][time_col])
856
+
857
+ data_test = data[data[time_col] < first_day_of_treatment].copy()
858
+ data_test[config.TG_COL] = 0
859
+ data_test.loc[data_test[unit_col].isin(treatment_group), config.TG_COL] = 1
860
+
861
+ if config.TIME_COUNTER_COL not in data_test.columns:
862
+ data_test = date_counter(
863
+ df = data_test,
864
+ date_col = time_col,
865
+ new_col = config.TIME_COUNTER_COL,
866
+ verbose = False
867
+ )
868
+ data_test[f"{config.TG_COL}_x_{config.TIME_COL}"] = data_test[config.TG_COL]*data_test[config.TIME_COUNTER_COL]
869
+
870
+ if len(data_test) > 0:
871
+
872
+ test_ols_model = ols(f'{outcome_col} ~ {config.TG_COL} + {config.TIME_COUNTER_COL} + {config.TG_COL}_x_{config.TIME_COL}', data = data_test).fit()
873
+ coef_TG_x_t_p = test_ols_model.pvalues[f"{config.TG_COL}_x_{config.TIME_COL}"]
888
874
 
889
- test_ols_model = ols(f'{outcome_col} ~ {config.TG_COL} + {config.TIME_COUNTER_COL} + {config.TG_COL}_x_{config.TIME_COL}', data = data_test).fit()
890
- coef_TG_x_t_p = test_ols_model.pvalues[f"{config.TG_COL}_x_{config.TIME_COL}"]
875
+ if coef_TG_x_t_p < alpha:
876
+ parallel = False
877
+ else:
878
+ parallel = True
891
879
 
892
- if coef_TG_x_t_p < alpha:
893
- parallel = False
880
+ else:
881
+ no_pre_period = True
882
+
894
883
  else:
895
- parallel = True
884
+ parallel = "not_tested"
885
+ test_ols_model = None
896
886
 
897
887
  else:
898
- parallel = "not_tested"
899
- test_ols_model = None
900
-
901
- if verbose:
902
- print("OK")
903
888
 
904
- if parallel == "not_tested":
905
- print("WARNING: Data could not be tested for parallel time trends.")
889
+ if len(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] == 1)]) > 0:
890
+
891
+ first_day_of_treatment = min(data[(data[unit_col].isin(treatment_group)) & (data[treatment_col] == 1)][time_col])
892
+
893
+ data_test = data[data[time_col] < first_day_of_treatment].copy()
894
+ data_test[config.TG_COL] = 0
895
+ data_test.loc[data_test[unit_col].isin(treatment_group), config.TG_COL] = 1
896
+
897
+ if config.TIME_COUNTER_COL not in data_test.columns:
898
+ data_test = date_counter(
899
+ df = data_test,
900
+ date_col = time_col,
901
+ new_col = config.TIME_COUNTER_COL,
902
+ verbose = False
903
+ )
904
+ data_test[f"{config.TG_COL}_x_{config.TIME_COL}"] = data_test[config.TG_COL]*data_test[config.TIME_COUNTER_COL]
905
+
906
+ if len(data_test) > 0:
907
+
908
+ test_ols_model = ols(f'{outcome_col} ~ {config.TG_COL} + {config.TIME_COUNTER_COL} + {config.TG_COL}_x_{config.TIME_COL}', data = data_test).fit()
909
+ coef_TG_x_t_p = test_ols_model.pvalues[f"{config.TG_COL}_x_{config.TIME_COL}"]
910
+
911
+ if coef_TG_x_t_p < alpha:
912
+ parallel = False
913
+ else:
914
+ parallel = True
915
+
916
+ else:
917
+ no_pre_period = True
918
+
919
+ else:
920
+ parallel = "not_tested"
921
+ test_ols_model = None
922
+
923
+ if verbose:
924
+ print("OK")
925
+
926
+ if not pre_post:
927
+
928
+ if parallel == "not_tested":
929
+ print("WARNING: Data could not be tested for parallel time trends.")
930
+ if no_pre_period:
931
+ print("WARNING: Data could not be tested for parallel time trends because there is no pre-treatment period.")
932
+
933
+ else:
934
+
935
+ if verbose:
936
+ print("NOTE: Data is pre-post data and parallel trends are not tested.")
906
937
 
907
938
  return [
908
939
  parallel,
@@ -1549,18 +1580,28 @@ def model_wrapper(
1549
1580
  model = SVR(kernel=svr_kernel)
1550
1581
 
1551
1582
  elif model_type == "xgb":
1552
- model = XGBRegressor(
1553
- learning_rate = xgb_learning_rate,
1554
- n_estimators = gb_iterations,
1555
- random_state = random_state
1556
- )
1583
+
1584
+ if XGBRegressor is not None:
1585
+ model = XGBRegressor(
1586
+ learning_rate = xgb_learning_rate,
1587
+ n_estimators = gb_iterations,
1588
+ random_state = random_state
1589
+ )
1590
+ else:
1591
+ model_estimation_error = True
1592
+ model_estimation_error_text = "XGBRegressor is not available. Please install xgboost to use this model type."
1557
1593
 
1558
1594
  elif model_type == "lgbm":
1559
- model = LGBMRegressor(
1560
- learning_rate = lgbm_learning_rate,
1561
- n_estimators = gb_iterations,
1562
- random_state = random_state
1563
- )
1595
+
1596
+ if LGBMRegressor is not None:
1597
+ model = LGBMRegressor(
1598
+ learning_rate = lgbm_learning_rate,
1599
+ n_estimators = gb_iterations,
1600
+ random_state = random_state
1601
+ )
1602
+ else:
1603
+ model_estimation_error = True
1604
+ model_estimation_error_text = "LGBMRegressor is not available. Please install lightgbm to use this model type."
1564
1605
 
1565
1606
  elif model_type == "mlp":
1566
1607
  model = Pipeline(
@@ -1,10 +1,11 @@
1
1
  Metadata-Version: 2.1
2
2
  Name: diffindiff
3
- Version: 2.5.5
3
+ Version: 2.5.6
4
4
  Summary: diffindiff: Python library for convenient Difference-in-Differences analyses
5
5
  Author: Thomas Wieland
6
6
  Author-email: geowieland@googlemail.com
7
7
  Description-Content-Type: text/markdown
8
+ Provides-Extra: optional
8
9
 
9
10
  # diffindiff: Python library for convenient Difference-in-Differences analyses
10
11
 
@@ -29,7 +30,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
29
30
 
30
31
  If you use this software, please cite:
31
32
 
32
- Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.5) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
33
+ Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.6) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
33
34
 
34
35
 
35
36
  ## Installation
@@ -177,11 +178,11 @@ See the /tests directory for usage examples of most of the included functions.
177
178
 
178
179
  ## AI Usage Statement
179
180
 
180
- This software was developed without the use of AI-generated code. The Continue Agent in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
181
+ This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
181
182
 
182
183
 
183
- ## What's new (v2.5.5)
184
+ ## What's new (v2.5.6)
184
185
 
185
186
  - Bugfixes
186
- - didanalysis.did_analysis(): Correct usage of param 'log_outcome' in any case and check whether value of 'log_outcome_add' is appropriate
187
- - didanalysis.DiffModel.plot() and didanalysis.DiffModel.plot_counterfactual(): Correct detection of the date format via object metadata
187
+ - didtools.is_parallel(): Checking whether there is a pre-treatment period and skipping test if not (additional NOTE)
188
+ - Optional installation of XGBoost and LightGBM (more stable if problems with the installation of these packages occur)
@@ -3,8 +3,10 @@ numpy
3
3
  statsmodels>=0.14.5
4
4
  scipy>=1.17
5
5
  scikit-learn
6
- xgboost
7
- lightgbm
8
6
  openpyxl
9
7
  matplotlib
10
8
  patsy
9
+
10
+ [optional]
11
+ lightgbm
12
+ xgboost
@@ -7,7 +7,7 @@ def read_README():
7
7
 
8
8
  setup(
9
9
  name='diffindiff',
10
- version='2.5.5',
10
+ version='2.5.6',
11
11
  description='diffindiff: Python library for convenient Difference-in-Differences analyses',
12
12
  packages=find_packages(include=["diffindiff", "diffindiff.tests"]),
13
13
  include_package_data=True,
@@ -25,11 +25,15 @@ setup(
25
25
  'statsmodels>=0.14.5',
26
26
  'scipy>=1.17',
27
27
  'scikit-learn',
28
- 'xgboost',
29
- 'lightgbm',
30
28
  'openpyxl',
31
29
  'matplotlib',
32
30
  'patsy',
33
31
  ],
32
+ extras_require={
33
+ "optional": [
34
+ 'lightgbm',
35
+ 'xgboost',
36
+ ]
37
+ },
34
38
  test_suite='tests',
35
39
  )
File without changes
File without changes