diffindiff 2.5.6__tar.gz → 2.5.7__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {diffindiff-2.5.6 → diffindiff-2.5.7}/PKG-INFO +24 -6
- {diffindiff-2.5.6 → diffindiff-2.5.7}/README.md +5 -4
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/config.py +3 -3
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/didanalysis_helper.py +70 -43
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/didtools.py +17 -3
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff.egg-info/PKG-INFO +24 -6
- {diffindiff-2.5.6 → diffindiff-2.5.7}/setup.py +1 -1
- {diffindiff-2.5.6 → diffindiff-2.5.7}/MANIFEST.in +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/__init__.py +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/didanalysis.py +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/diddata.py +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/tests/__init__.py +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/tests/data/Corona_Hesse.xlsx +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/tests/data/counties_DE.csv +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/tests/data/curfew_DE.csv +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff/tests/tests_diffindiff.py +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff.egg-info/SOURCES.txt +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff.egg-info/dependency_links.txt +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff.egg-info/requires.txt +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/diffindiff.egg-info/top_level.txt +0 -0
- {diffindiff-2.5.6 → diffindiff-2.5.7}/setup.cfg +0 -0
|
@@ -1,11 +1,28 @@
|
|
|
1
|
-
Metadata-Version: 2.
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
2
|
Name: diffindiff
|
|
3
|
-
Version: 2.5.
|
|
3
|
+
Version: 2.5.7
|
|
4
4
|
Summary: diffindiff: Python library for convenient Difference-in-Differences analyses
|
|
5
5
|
Author: Thomas Wieland
|
|
6
6
|
Author-email: geowieland@googlemail.com
|
|
7
7
|
Description-Content-Type: text/markdown
|
|
8
|
+
Requires-Dist: pandas
|
|
9
|
+
Requires-Dist: numpy
|
|
10
|
+
Requires-Dist: statsmodels>=0.14.5
|
|
11
|
+
Requires-Dist: scipy>=1.17
|
|
12
|
+
Requires-Dist: scikit-learn
|
|
13
|
+
Requires-Dist: openpyxl
|
|
14
|
+
Requires-Dist: matplotlib
|
|
15
|
+
Requires-Dist: patsy
|
|
8
16
|
Provides-Extra: optional
|
|
17
|
+
Requires-Dist: lightgbm; extra == "optional"
|
|
18
|
+
Requires-Dist: xgboost; extra == "optional"
|
|
19
|
+
Dynamic: author
|
|
20
|
+
Dynamic: author-email
|
|
21
|
+
Dynamic: description
|
|
22
|
+
Dynamic: description-content-type
|
|
23
|
+
Dynamic: provides-extra
|
|
24
|
+
Dynamic: requires-dist
|
|
25
|
+
Dynamic: summary
|
|
9
26
|
|
|
10
27
|
# diffindiff: Python library for convenient Difference-in-Differences analyses
|
|
11
28
|
|
|
@@ -30,7 +47,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
|
|
|
30
47
|
|
|
31
48
|
If you use this software, please cite:
|
|
32
49
|
|
|
33
|
-
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.
|
|
50
|
+
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.7) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
|
|
34
51
|
|
|
35
52
|
|
|
36
53
|
## Installation
|
|
@@ -181,8 +198,9 @@ See the /tests directory for usage examples of most of the included functions.
|
|
|
181
198
|
This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
|
|
182
199
|
|
|
183
200
|
|
|
184
|
-
## What's new (v2.5.
|
|
201
|
+
## What's new (v2.5.7)
|
|
185
202
|
|
|
186
203
|
- Bugfixes
|
|
187
|
-
-
|
|
188
|
-
-
|
|
204
|
+
- Catching RecursionError when model formulas include too many variables
|
|
205
|
+
- Checking input data frames for duplicated columns
|
|
206
|
+
- Avoiding treatment group deviation in summary when treatment/control groups are identical
|
|
@@ -21,7 +21,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
|
|
|
21
21
|
|
|
22
22
|
If you use this software, please cite:
|
|
23
23
|
|
|
24
|
-
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.
|
|
24
|
+
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.7) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
|
|
25
25
|
|
|
26
26
|
|
|
27
27
|
## Installation
|
|
@@ -172,8 +172,9 @@ See the /tests directory for usage examples of most of the included functions.
|
|
|
172
172
|
This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
|
|
173
173
|
|
|
174
174
|
|
|
175
|
-
## What's new (v2.5.
|
|
175
|
+
## What's new (v2.5.7)
|
|
176
176
|
|
|
177
177
|
- Bugfixes
|
|
178
|
-
-
|
|
179
|
-
-
|
|
178
|
+
- Catching RecursionError when model formulas include too many variables
|
|
179
|
+
- Checking input data frames for duplicated columns
|
|
180
|
+
- Avoiding treatment group deviation in summary when treatment/control groups are identical
|
|
@@ -4,15 +4,15 @@
|
|
|
4
4
|
# Author: Thomas Wieland
|
|
5
5
|
# ORCID: 0000-0001-5168-9846
|
|
6
6
|
# mail: geowieland@googlemail.com
|
|
7
|
-
# Version: 1.0.
|
|
8
|
-
# Last update: 2026-10-
|
|
7
|
+
# Version: 1.0.28
|
|
8
|
+
# Last update: 2026-10-10 09:58
|
|
9
9
|
# Copyright (c) 2025-2026 Thomas Wieland
|
|
10
10
|
#-----------------------------------------------------------------------
|
|
11
11
|
|
|
12
12
|
# Basic config:
|
|
13
13
|
|
|
14
14
|
PACKAGE_NAME = "diffindiff"
|
|
15
|
-
PACKAGE_VERSION = "2.5.
|
|
15
|
+
PACKAGE_VERSION = "2.5.7"
|
|
16
16
|
|
|
17
17
|
VERBOSE = False
|
|
18
18
|
|
|
@@ -4,8 +4,8 @@
|
|
|
4
4
|
# Author: Thomas Wieland
|
|
5
5
|
# ORCID: 0000-0001-5168-9846
|
|
6
6
|
# mail: geowieland@googlemail.com
|
|
7
|
-
# Version: 1.2.
|
|
8
|
-
# Last update: 2026-
|
|
7
|
+
# Version: 1.2.5
|
|
8
|
+
# Last update: 2026-10-09 23:04
|
|
9
9
|
# Copyright (c) 2025-2026 Thomas Wieland
|
|
10
10
|
#-----------------------------------------------------------------------
|
|
11
11
|
|
|
@@ -1085,33 +1085,47 @@ def ols_fit(
|
|
|
1085
1085
|
>>> ols_fit(df, 'y ~ x')
|
|
1086
1086
|
"""
|
|
1087
1087
|
|
|
1088
|
+
ols_model = None
|
|
1089
|
+
ols_coef = None
|
|
1090
|
+
ols_coef_se = None
|
|
1091
|
+
ols_coef_t = None
|
|
1092
|
+
ols_coef_p = None
|
|
1093
|
+
ols_coef_ci = None
|
|
1094
|
+
ols_predictions = None
|
|
1095
|
+
|
|
1088
1096
|
if verbose:
|
|
1089
1097
|
print("Estimating model via Ordinary Least Squares", end = " ... ")
|
|
1090
1098
|
|
|
1091
|
-
|
|
1099
|
+
try:
|
|
1092
1100
|
|
|
1093
|
-
|
|
1094
|
-
|
|
1095
|
-
|
|
1096
|
-
|
|
1097
|
-
|
|
1098
|
-
|
|
1099
|
-
|
|
1101
|
+
if cluster_SE_by is not None:
|
|
1102
|
+
|
|
1103
|
+
ols_model = ols(
|
|
1104
|
+
formula,
|
|
1105
|
+
data=data
|
|
1106
|
+
).fit(
|
|
1107
|
+
cov_type="cluster",
|
|
1108
|
+
cov_kwds={"groups": data[cluster_SE_by]} if cluster_SE_by else None
|
|
1109
|
+
)
|
|
1100
1110
|
|
|
1101
|
-
|
|
1111
|
+
else:
|
|
1112
|
+
|
|
1113
|
+
ols_model = ols(formula, data).fit()
|
|
1114
|
+
|
|
1115
|
+
ols_coef = ols_model.params
|
|
1116
|
+
ols_coef_se = ols_model.bse
|
|
1117
|
+
ols_coef_t = ols_model.tvalues
|
|
1118
|
+
ols_coef_p = ols_model.pvalues
|
|
1119
|
+
ols_coef_ci = ols_model.conf_int(alpha = confint_alpha)
|
|
1102
1120
|
|
|
1103
|
-
|
|
1121
|
+
ols_predictions = ols_model.get_prediction(data).summary_frame(alpha=confint_alpha)
|
|
1104
1122
|
|
|
1105
|
-
|
|
1106
|
-
|
|
1107
|
-
|
|
1108
|
-
|
|
1109
|
-
|
|
1110
|
-
|
|
1111
|
-
ols_predictions = ols_model.get_prediction(data).summary_frame(alpha=confint_alpha)
|
|
1112
|
-
|
|
1113
|
-
if verbose:
|
|
1114
|
-
print("OK")
|
|
1123
|
+
if verbose:
|
|
1124
|
+
print("OK")
|
|
1125
|
+
|
|
1126
|
+
except RecursionError as e:
|
|
1127
|
+
|
|
1128
|
+
raise ValueError(f"OLS estimation has induced a RecursionError: {str(e)}. There are likely too many variables in the model. Try sys.setrecursionlimit() or diff-in-diff analysis with demean = True")
|
|
1115
1129
|
|
|
1116
1130
|
return [
|
|
1117
1131
|
ols_model,
|
|
@@ -1160,26 +1174,32 @@ def ml_fit(
|
|
|
1160
1174
|
|
|
1161
1175
|
if verbose:
|
|
1162
1176
|
print("Estimating model via Maximum Likelihood", end = " ... ")
|
|
1163
|
-
|
|
1164
|
-
y, X = dmatrices(
|
|
1165
|
-
formula,
|
|
1166
|
-
data=data,
|
|
1167
|
-
return_type = "dataframe"
|
|
1168
|
-
)
|
|
1169
1177
|
|
|
1170
|
-
|
|
1171
|
-
y,
|
|
1172
|
-
X,
|
|
1173
|
-
family=family(link=link)
|
|
1174
|
-
).fit()
|
|
1175
|
-
|
|
1176
|
-
mle_coef = mle_model.params
|
|
1177
|
-
mle_coef_se = mle_model.bse
|
|
1178
|
-
mle_coef_z = mle_model.tvalues
|
|
1179
|
-
mle_coef_p = mle_model.pvalues
|
|
1180
|
-
mle_coef_ci = mle_model.conf_int(alpha=confint_alpha)
|
|
1178
|
+
try:
|
|
1181
1179
|
|
|
1182
|
-
|
|
1180
|
+
y, X = dmatrices(
|
|
1181
|
+
formula,
|
|
1182
|
+
data=data,
|
|
1183
|
+
return_type = "dataframe"
|
|
1184
|
+
)
|
|
1185
|
+
|
|
1186
|
+
mle_model = sm.GLM(
|
|
1187
|
+
y,
|
|
1188
|
+
X,
|
|
1189
|
+
family=family(link=link)
|
|
1190
|
+
).fit()
|
|
1191
|
+
|
|
1192
|
+
mle_coef = mle_model.params
|
|
1193
|
+
mle_coef_se = mle_model.bse
|
|
1194
|
+
mle_coef_z = mle_model.tvalues
|
|
1195
|
+
mle_coef_p = mle_model.pvalues
|
|
1196
|
+
mle_coef_ci = mle_model.conf_int(alpha=confint_alpha)
|
|
1197
|
+
|
|
1198
|
+
mle_predictions = mle_model.get_prediction(X).summary_frame(alpha=confint_alpha)
|
|
1199
|
+
|
|
1200
|
+
except RecursionError as e:
|
|
1201
|
+
|
|
1202
|
+
raise ValueError(f"ML estimation has induced a RecursionError: {str(e)}. There are likely too many variables in the model. Try sys.setrecursionlimit() or diff-in-diff analysis with demean = True")
|
|
1183
1203
|
|
|
1184
1204
|
if verbose:
|
|
1185
1205
|
print("OK")
|
|
@@ -1333,7 +1353,7 @@ def extract_model_results(
|
|
|
1333
1353
|
}
|
|
1334
1354
|
|
|
1335
1355
|
if config.REMOVE_DUPLICATES_FROM_RESULTS_DICT:
|
|
1336
|
-
beta_1 = remove_duplicates_from_dict(beta_1)
|
|
1356
|
+
beta_1 = remove_duplicates_from_dict(beta_1, ignore_key=config.OLS_MODEL_RESULTS["coef_name"]["model_results_key"])
|
|
1337
1357
|
|
|
1338
1358
|
model_results[config.EFFECTS_TYPES["beta_1"]["model_results_key"]] = beta_1
|
|
1339
1359
|
|
|
@@ -1800,7 +1820,8 @@ def create_timestamp(function) -> dict:
|
|
|
1800
1820
|
return timestamp_dict
|
|
1801
1821
|
|
|
1802
1822
|
def remove_duplicates_from_dict(
|
|
1803
|
-
any_dict: dict
|
|
1823
|
+
any_dict: dict,
|
|
1824
|
+
ignore_key = None
|
|
1804
1825
|
) -> dict:
|
|
1805
1826
|
|
|
1806
1827
|
"""
|
|
@@ -1822,7 +1843,13 @@ def remove_duplicates_from_dict(
|
|
|
1822
1843
|
|
|
1823
1844
|
for key, value in any_dict.items():
|
|
1824
1845
|
|
|
1825
|
-
|
|
1846
|
+
if ignore_key is not None:
|
|
1847
|
+
identifier = tuple(sorted(
|
|
1848
|
+
(k, v) for k, v in value.items()
|
|
1849
|
+
if k != ignore_key
|
|
1850
|
+
))
|
|
1851
|
+
else:
|
|
1852
|
+
identifier = tuple(sorted(value.items()))
|
|
1826
1853
|
|
|
1827
1854
|
if identifier not in seen:
|
|
1828
1855
|
seen.add(identifier)
|
|
@@ -4,8 +4,8 @@
|
|
|
4
4
|
# Author: Thomas Wieland
|
|
5
5
|
# ORCID: 0000-0001-5168-9846
|
|
6
6
|
# mail: geowieland@googlemail.com
|
|
7
|
-
# Version: 2.2.
|
|
8
|
-
# Last update: 2026-10-
|
|
7
|
+
# Version: 2.2.7
|
|
8
|
+
# Last update: 2026-10-10 09:57
|
|
9
9
|
# Copyright (c) 2025-2026 Thomas Wieland
|
|
10
10
|
#-----------------------------------------------------------------------
|
|
11
11
|
|
|
@@ -46,7 +46,8 @@ def check_columns(
|
|
|
46
46
|
):
|
|
47
47
|
|
|
48
48
|
"""
|
|
49
|
-
Check that the given columns exist in a DataFrame
|
|
49
|
+
Check that the given columns exist in a DataFrame
|
|
50
|
+
and whether there are duplicated column names.
|
|
50
51
|
|
|
51
52
|
Parameters
|
|
52
53
|
----------
|
|
@@ -65,6 +66,8 @@ def check_columns(
|
|
|
65
66
|
------
|
|
66
67
|
KeyError
|
|
67
68
|
If any column from ``columns`` is missing in ``df``.
|
|
69
|
+
KeyError
|
|
70
|
+
If any column from ``columns`` is duplicated.
|
|
68
71
|
|
|
69
72
|
Examples
|
|
70
73
|
--------
|
|
@@ -86,6 +89,17 @@ def check_columns(
|
|
|
86
89
|
if missing_columns:
|
|
87
90
|
raise KeyError(f"Data do not contain column(s): {', '.join(missing_columns)}")
|
|
88
91
|
|
|
92
|
+
if verbose:
|
|
93
|
+
print("Checking whether columns are duplicated in data frame", end = " ... ")
|
|
94
|
+
|
|
95
|
+
cols_duplicated = df.columns[df.columns.duplicated()].tolist()
|
|
96
|
+
|
|
97
|
+
if verbose:
|
|
98
|
+
print("OK")
|
|
99
|
+
|
|
100
|
+
if len(cols_duplicated) > 0:
|
|
101
|
+
raise KeyError(f"Data contain duplicated relevant columns: {', '.join(cols_duplicated)}")
|
|
102
|
+
|
|
89
103
|
def is_numeric(
|
|
90
104
|
df: pd.DataFrame,
|
|
91
105
|
columns: list,
|
|
@@ -1,11 +1,28 @@
|
|
|
1
|
-
Metadata-Version: 2.
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
2
|
Name: diffindiff
|
|
3
|
-
Version: 2.5.
|
|
3
|
+
Version: 2.5.7
|
|
4
4
|
Summary: diffindiff: Python library for convenient Difference-in-Differences analyses
|
|
5
5
|
Author: Thomas Wieland
|
|
6
6
|
Author-email: geowieland@googlemail.com
|
|
7
7
|
Description-Content-Type: text/markdown
|
|
8
|
+
Requires-Dist: pandas
|
|
9
|
+
Requires-Dist: numpy
|
|
10
|
+
Requires-Dist: statsmodels>=0.14.5
|
|
11
|
+
Requires-Dist: scipy>=1.17
|
|
12
|
+
Requires-Dist: scikit-learn
|
|
13
|
+
Requires-Dist: openpyxl
|
|
14
|
+
Requires-Dist: matplotlib
|
|
15
|
+
Requires-Dist: patsy
|
|
8
16
|
Provides-Extra: optional
|
|
17
|
+
Requires-Dist: lightgbm; extra == "optional"
|
|
18
|
+
Requires-Dist: xgboost; extra == "optional"
|
|
19
|
+
Dynamic: author
|
|
20
|
+
Dynamic: author-email
|
|
21
|
+
Dynamic: description
|
|
22
|
+
Dynamic: description-content-type
|
|
23
|
+
Dynamic: provides-extra
|
|
24
|
+
Dynamic: requires-dist
|
|
25
|
+
Dynamic: summary
|
|
9
26
|
|
|
10
27
|
# diffindiff: Python library for convenient Difference-in-Differences analyses
|
|
11
28
|
|
|
@@ -30,7 +47,7 @@ A case study that utilizes the diffindiff library is available on [arXiv](https:
|
|
|
30
47
|
|
|
31
48
|
If you use this software, please cite:
|
|
32
49
|
|
|
33
|
-
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.
|
|
50
|
+
Wieland, T. (2026). diffindiff: A Python library for convenient difference-in-differences analyses (Version 2.5.7) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.18656820
|
|
34
51
|
|
|
35
52
|
|
|
36
53
|
## Installation
|
|
@@ -181,8 +198,9 @@ See the /tests directory for usage examples of most of the included functions.
|
|
|
181
198
|
This software was developed without the use of AI-generated code. The GitHub Copilot Chat in Microsoft Visual Studio Code using the GPT-5 mini model (by OpenAI) was used solely to assist in drafting and refining docstrings for documentation. The corresponding guidelines and constraints defined by the author are documented in `AGENTS-docstrings.md` in the [public GitHub repository](https://github.com/geowieland/diffindiff_official).
|
|
182
199
|
|
|
183
200
|
|
|
184
|
-
## What's new (v2.5.
|
|
201
|
+
## What's new (v2.5.7)
|
|
185
202
|
|
|
186
203
|
- Bugfixes
|
|
187
|
-
-
|
|
188
|
-
-
|
|
204
|
+
- Catching RecursionError when model formulas include too many variables
|
|
205
|
+
- Checking input data frames for duplicated columns
|
|
206
|
+
- Avoiding treatment group deviation in summary when treatment/control groups are identical
|
|
@@ -7,7 +7,7 @@ def read_README():
|
|
|
7
7
|
|
|
8
8
|
setup(
|
|
9
9
|
name='diffindiff',
|
|
10
|
-
version='2.5.
|
|
10
|
+
version='2.5.7',
|
|
11
11
|
description='diffindiff: Python library for convenient Difference-in-Differences analyses',
|
|
12
12
|
packages=find_packages(include=["diffindiff", "diffindiff.tests"]),
|
|
13
13
|
include_package_data=True,
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|