seastats 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
seastats-0.1.0/LICENSE ADDED
@@ -0,0 +1,274 @@
1
+ European Union Public Licence
2
+ V. 1.2
3
+
4
+ EUPL © the European Union 2007, 2016
5
+
6
+ This European Union Public Licence (the ‘EUPL’) applies to the Work (as
7
+ defined below) which is provided under the terms of this Licence. Any use of
8
+ the Work, other than as authorised under this Licence is prohibited (to the
9
+ extent such use is covered by a right of the copyright holder of the Work).
10
+
11
+ The Work is provided under the terms of this Licence when the Licensor (as
12
+ defined below) has placed the following notice immediately following the
13
+ copyright notice for the Work: “Licensed under the EUPL”, or has expressed by
14
+ any other means his willingness to license under the EUPL.
15
+
16
+ 1. Definitions
17
+
18
+ In this Licence, the following terms have the following meaning:
19
+ — ‘The Licence’: this Licence.
20
+ — ‘The Original Work’: the work or software distributed or communicated by the
21
+ ‘Licensor under this Licence, available as Source Code and also as
22
+ ‘Executable Code as the case may be.
23
+ — ‘Derivative Works’: the works or software that could be created by the
24
+ ‘Licensee, based upon the Original Work or modifications thereof. This
25
+ ‘Licence does not define the extent of modification or dependence on the
26
+ ‘Original Work required in order to classify a work as a Derivative Work;
27
+ ‘this extent is determined by copyright law applicable in the country
28
+ ‘mentioned in Article 15.
29
+ — ‘The Work’: the Original Work or its Derivative Works.
30
+ — ‘The Source Code’: the human-readable form of the Work which is the most
31
+ convenient for people to study and modify.
32
+
33
+ — ‘The Executable Code’: any code which has generally been compiled and which
34
+ is meant to be interpreted by a computer as a program.
35
+ — ‘The Licensor’: the natural or legal person that distributes or communicates
36
+ the Work under the Licence.
37
+ — ‘Contributor(s)’: any natural or legal person who modifies the Work under
38
+ the Licence, or otherwise contributes to the creation of a Derivative Work.
39
+ — ‘The Licensee’ or ‘You’: any natural or legal person who makes any usage of
40
+ the Work under the terms of the Licence.
41
+ — ‘Distribution’ or ‘Communication’: any act of selling, giving, lending,
42
+ renting, distributing, communicating, transmitting, or otherwise making
43
+ available, online or offline, copies of the Work or providing access to its
44
+ essential functionalities at the disposal of any other natural or legal
45
+ person.
46
+
47
+ 2. Scope of the rights granted by the Licence
48
+
49
+ The Licensor hereby grants You a worldwide, royalty-free, non-exclusive,
50
+ sublicensable licence to do the following, for the duration of copyright
51
+ vested in the Original Work:
52
+
53
+ — use the Work in any circumstance and for all usage,
54
+ — reproduce the Work,
55
+ — modify the Work, and make Derivative Works based upon the Work,
56
+ — communicate to the public, including the right to make available or display
57
+ the Work or copies thereof to the public and perform publicly, as the case
58
+ may be, the Work,
59
+ — distribute the Work or copies thereof,
60
+ — lend and rent the Work or copies thereof,
61
+ — sublicense rights in the Work or copies thereof.
62
+
63
+ Those rights can be exercised on any media, supports and formats, whether now
64
+ known or later invented, as far as the applicable law permits so.
65
+
66
+ In the countries where moral rights apply, the Licensor waives his right to
67
+ exercise his moral right to the extent allowed by law in order to make
68
+ effective the licence of the economic rights here above listed.
69
+
70
+ The Licensor grants to the Licensee royalty-free, non-exclusive usage rights
71
+ to any patents held by the Licensor, to the extent necessary to make use of
72
+ the rights granted on the Work under this Licence.
73
+
74
+ 3. Communication of the Source Code
75
+
76
+ The Licensor may provide the Work either in its Source Code form, or as
77
+ Executable Code. If the Work is provided as Executable Code, the Licensor
78
+ provides in addition a machine-readable copy of the Source Code of the Work
79
+ along with each copy of the Work that the Licensor distributes or indicates,
80
+ in a notice following the copyright notice attached to the Work, a repository
81
+ where the Source Code is easily and freely accessible for as long as the
82
+ Licensor continues to distribute or communicate the Work.
83
+
84
+ 4. Limitations on copyright
85
+
86
+ Nothing in this Licence is intended to deprive the Licensee of the benefits
87
+ from any exception or limitation to the exclusive rights of the rights owners
88
+ in the Work, of the exhaustion of those rights or of other applicable
89
+ limitations thereto.
90
+
91
+ 5. Obligations of the Licensee
92
+
93
+ The grant of the rights mentioned above is subject to some restrictions and
94
+ obligations imposed on the Licensee. Those obligations are the following:
95
+
96
+ Attribution right: The Licensee shall keep intact all copyright, patent or
97
+ trademarks notices and all notices that refer to the Licence and to the
98
+ disclaimer of warranties. The Licensee must include a copy of such notices and
99
+ a copy of the Licence with every copy of the Work he/she distributes or
100
+ communicates. The Licensee must cause any Derivative Work to carry prominent
101
+ notices stating that the Work has been modified and the date of modification.
102
+
103
+ Copyleft clause: If the Licensee distributes or communicates copies of the
104
+ Original Works or Derivative Works, this Distribution or Communication will be
105
+ done under the terms of this Licence or of a later version of this Licence
106
+ unless the Original Work is expressly distributed only under this version of
107
+ the Licence — for example by communicating ‘EUPL v. 1.2 only’. The Licensee
108
+ (becoming Licensor) cannot offer or impose any additional terms or conditions
109
+ on the Work or Derivative Work that alter or restrict the terms of the
110
+ Licence.
111
+
112
+ Compatibility clause: If the Licensee Distributes or Communicates Derivative
113
+ Works or copies thereof based upon both the Work and another work licensed
114
+ under a Compatible Licence, this Distribution or Communication can be done
115
+ under the terms of this Compatible Licence. For the sake of this clause,
116
+ ‘Compatible Licence’ refers to the licences listed in the appendix attached to
117
+ this Licence. Should the Licensee's obligations under the Compatible Licence
118
+ conflict with his/her obligations under this Licence, the obligations of the
119
+ Compatible Licence shall prevail.
120
+
121
+ Provision of Source Code: When distributing or communicating copies of the
122
+ Work, the Licensee will provide a machine-readable copy of the Source Code or
123
+ indicate a repository where this Source will be easily and freely available
124
+ for as long as the Licensee continues to distribute or communicate the Work.
125
+
126
+ Legal Protection: This Licence does not grant permission to use the trade
127
+ names, trademarks, service marks, or names of the Licensor, except as required
128
+ for reasonable and customary use in describing the origin of the Work and
129
+ reproducing the content of the copyright notice.
130
+
131
+ 6. Chain of Authorship
132
+
133
+ The original Licensor warrants that the copyright in the Original Work granted
134
+ hereunder is owned by him/her or licensed to him/her and that he/she has the
135
+ power and authority to grant the Licence.
136
+
137
+ Each Contributor warrants that the copyright in the modifications he/she
138
+ brings to the Work are owned by him/her or licensed to him/her and that he/she
139
+ has the power and authority to grant the Licence.
140
+
141
+ Each time You accept the Licence, the original Licensor and subsequent
142
+ Contributors grant You a licence to their contributions to the Work, under the
143
+ terms of this Licence.
144
+
145
+ 7. Disclaimer of Warranty
146
+
147
+ The Work is a work in progress, which is continuously improved by numerous
148
+ Contributors. It is not a finished work and may therefore contain defects or
149
+ ‘bugs’ inherent to this type of development.
150
+
151
+ For the above reason, the Work is provided under the Licence on an ‘as is’
152
+ basis and without warranties of any kind concerning the Work, including
153
+ without limitation merchantability, fitness for a particular purpose, absence
154
+ of defects or errors, accuracy, non-infringement of intellectual property
155
+ rights other than copyright as stated in Article 6 of this Licence.
156
+
157
+ This disclaimer of warranty is an essential part of the Licence and a
158
+ condition for the grant of any rights to the Work.
159
+
160
+ 8. Disclaimer of Liability
161
+
162
+ Except in the cases of wilful misconduct or damages directly caused to natural
163
+ persons, the Licensor will in no event be liable for any direct or indirect,
164
+ material or moral, damages of any kind, arising out of the Licence or of the
165
+ use of the Work, including without limitation, damages for loss of goodwill,
166
+ work stoppage, computer failure or malfunction, loss of data or any commercial
167
+ damage, even if the Licensor has been advised of the possibility of such
168
+ damage. However, the Licensor will be liable under statutory product liability
169
+ laws as far such laws apply to the Work.
170
+
171
+ 9. Additional agreements
172
+
173
+ While distributing the Work, You may choose to conclude an additional
174
+ agreement, defining obligations or services consistent with this Licence.
175
+ However, if accepting obligations, You may act only on your own behalf and on
176
+ your sole responsibility, not on behalf of the original Licensor or any other
177
+ Contributor, and only if You agree to indemnify, defend, and hold each
178
+ Contributor harmless for any liability incurred by, or claims asserted against
179
+ such Contributor by the fact You have accepted any warranty or additional
180
+ liability.
181
+
182
+ 10. Acceptance of the Licence
183
+
184
+ The provisions of this Licence can be accepted by clicking on an icon ‘I
185
+ agree’ placed under the bottom of a window displaying the text of this Licence
186
+ or by affirming consent in any other similar way, in accordance with the rules
187
+ of applicable law. Clicking on that icon indicates your clear and irrevocable
188
+ acceptance of this Licence and all of its terms and conditions.
189
+
190
+ Similarly, you irrevocably accept this Licence and all of its terms and
191
+ conditions by exercising any rights granted to You by Article 2 of this
192
+ Licence, such as the use of the Work, the creation by You of a Derivative Work
193
+ or the Distribution or Communication by You of the Work or copies thereof.
194
+
195
+ 11. Information to the public
196
+
197
+ In case of any Distribution or Communication of the Work by means of
198
+ electronic communication by You (for example, by offering to download the Work
199
+ from a remote location) the distribution channel or media (for example, a
200
+ website) must at least provide to the public the information requested by the
201
+ applicable law regarding the Licensor, the Licence and the way it may be
202
+ accessible, concluded, stored and reproduced by the Licensee.
203
+
204
+ 12. Termination of the Licence
205
+
206
+ The Licence and the rights granted hereunder will terminate automatically upon
207
+ any breach by the Licensee of the terms of the Licence. Such a termination
208
+ will not terminate the licences of any person who has received the Work from
209
+ the Licensee under the Licence, provided such persons remain in full
210
+ compliance with the Licence.
211
+
212
+ 13. Miscellaneous
213
+
214
+ Without prejudice of Article 9 above, the Licence represents the complete
215
+ agreement between the Parties as to the Work.
216
+
217
+ If any provision of the Licence is invalid or unenforceable under applicable
218
+ law, this will not affect the validity or enforceability of the Licence as a
219
+ whole. Such provision will be construed or reformed so as necessary to make it
220
+ valid and enforceable.
221
+
222
+ The European Commission may publish other linguistic versions or new versions
223
+ of this Licence or updated versions of the Appendix, so far this is required
224
+ and reasonable, without reducing the scope of the rights granted by the
225
+ Licence. New versions of the Licence will be published with a unique version
226
+ number.
227
+
228
+ All linguistic versions of this Licence, approved by the European Commission,
229
+ have identical value. Parties can take advantage of the linguistic version of
230
+ their choice.
231
+
232
+ 14. Jurisdiction
233
+
234
+ Without prejudice to specific agreement between parties,
235
+ — any litigation resulting from the interpretation of this License, arising
236
+ between the European Union institutions, bodies, offices or agencies, as a
237
+ Licensor, and any Licensee, will be subject to the jurisdiction of the Court
238
+ of Justice of the European Union, as laid down in article 272 of the Treaty
239
+ on the Functioning of the European Union,
240
+ — any litigation arising between other parties and resulting from the
241
+ interpretation of this License, will be subject to the exclusive
242
+ jurisdiction of the competent court where the Licensor resides or conducts
243
+ its primary business.
244
+
245
+ 15. Applicable Law
246
+
247
+ Without prejudice to specific agreement between parties,
248
+ — this Licence shall be governed by the law of the European Union Member State
249
+ where the Licensor has his seat, resides or has his registered office,
250
+ — this licence shall be governed by Belgian law if the Licensor has no seat,
251
+ residence or registered office inside a European Union Member State.
252
+
253
+ Appendix
254
+
255
+ ‘Compatible Licences’ according to Article 5 EUPL are:
256
+ — GNU General Public License (GPL) v. 2, v. 3
257
+ — GNU Affero General Public License (AGPL) v. 3
258
+ — Open Software License (OSL) v. 2.1, v. 3.0
259
+ — Eclipse Public License (EPL) v. 1.0
260
+ — CeCILL v. 2.0, v. 2.1
261
+ — Mozilla Public Licence (MPL) v. 2
262
+ — GNU Lesser General Public Licence (LGPL) v. 2.1, v. 3
263
+ — Creative Commons Attribution-ShareAlike v. 3.0 Unported (CC BY-SA 3.0) for
264
+ works other than software
265
+ — European Union Public Licence (EUPL) v. 1.1, v. 1.2
266
+ — Québec Free and Open-Source Licence — Reciprocity (LiLiQ-R) or
267
+ Strong Reciprocity (LiLiQ-R+)
268
+
269
+ — The European Commission may update this Appendix to later versions of the
270
+ above licences without producing a new version of the EUPL, as long as they
271
+ provide the rights granted in Article 2 of this Licence and protect the
272
+ covered Source Code from exclusive appropriation.
273
+ — All other changes or additions to this Appendix require the production of a
274
+ new EUPL version.
@@ -0,0 +1,171 @@
1
+ Metadata-Version: 2.1
2
+ Name: seastats
3
+ Version: 0.1.0
4
+ Summary: package for metocean statistics
5
+ Home-page: https://github.com/oceanmodeling/seastats.git
6
+ License: EUPL-1.2
7
+ Author: tomsail
8
+ Author-email: saillour.thomas@gmail.com
9
+ Requires-Python: >=3.10,<4.0
10
+ Classifier: Development Status :: 4 - Beta
11
+ Classifier: Environment :: Other Environment
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Intended Audience :: Science/Research
14
+ Classifier: License :: OSI Approved :: European Union Public Licence 1.2 (EUPL 1.2)
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3.10
19
+ Classifier: Programming Language :: Python :: 3.11
20
+ Classifier: Programming Language :: Python :: 3.12
21
+ Classifier: Programming Language :: Python :: 3.13
22
+ Classifier: Topic :: Scientific/Engineering :: Physics
23
+ Requires-Dist: numpy
24
+ Requires-Dist: pandas[performance,pyarrow]
25
+ Requires-Dist: pyextremes
26
+ Project-URL: Repository, https://github.com/oceanmodeling/seastats.git
27
+ Description-Content-Type: text/markdown
28
+
29
+ # SeaStats
30
+
31
+ `seastats` is a simple package to compare and analyse 2 time series. We use the following convention in this repo:
32
+ * `sim`: modelled surge time series
33
+ * `mod`: observed surge time series
34
+
35
+ The main function is:
36
+
37
+ ```python
38
+ def get_stats(
39
+ sim: Series,
40
+ obs: Series,
41
+ metrics: Sequence[str] = SUGGESTED_METRICS,
42
+ quantile: float = 0,
43
+ cluster: int = 72,
44
+ round: int = -1
45
+ ) -> dict[str, float]
46
+ ```
47
+ Calculates various statistical metrics between the simulated and observed time series data.
48
+ ## Parameters:
49
+ * **sim** (pd.Series). The simulated time series data.
50
+ * **obs** (pd.Series). The observed time series data.
51
+ * **metrics** (list[str]). (Optional) The list of statistical metrics to calculate. If metrics = ["all"], all items in `SUPPORTED_METRICS` will be calculated. Default is all items in `SUGGESTED_METRICS`.
52
+ * **quantile** (float). (Optional) Quantile used to calculate the metrics. Default is `0` (no selection)
53
+ * **cluster** (int). (Optional) Cluster duration for grouping storm events. Default is `72` hours.
54
+ * **round** (int). (Optional) Apply rounding to the results to. Default is no rounding (value is `-1`)
55
+
56
+ Returns a dictionary containing the calculated metrics and their corresponding values. With 2 types of metrics:
57
+ * [The "general" metrics](#general-metrics): All the basic metrics needed for signal comparison (RMSE, RMS, Correlation etc..). See details below
58
+ * `bias`: Bias
59
+ * `rmse`: Root Mean Square Error
60
+ * `rms`: Root Mean Square
61
+ * `rms_95`: Root Mean Square for data points above 95th percentile
62
+ * `sim_mean`: Mean of simulated values
63
+ * `obs_mean`: Mean of observed values
64
+ * `sim_std`: Standard deviation of simulated values
65
+ * `obs_std`: Standard deviation of observed values
66
+ * `mae`: Mean Absolute Error
67
+ * `mse`: Mean Square Error
68
+ * `nse`: Nash-Sutcliffe Efficiency
69
+ * `lamba`: Lambda index
70
+ * `cr`: Pearson Correlation coefficient
71
+ * `cr_95`: Pearson Correlation coefficient for data points above 95th percentile
72
+ * `slope`: Slope of Model/Obs correlation
73
+ * `intercept`: Intercept of Model/Obs correlation
74
+ * `slope_pp`: Slope of Model/Obs correlation of percentiles
75
+ * `intercept_pp`: Intercept of Model/Obs correlation of percentiles
76
+ * `mad`: Mean Absolute Deviation
77
+ * `madp`: Mean Absolute Deviation of percentiles
78
+ * `madc`: `mad + madp`
79
+ * `kge`: Kling–Gupta Efficiency
80
+ * [The storm metrics](#storm-metrics): a PoT selection is done on the observed signal (using the `match_extremes()` function). Function returns the decreasing extreme event peak values for observed and modeled signals (and time lag between events). See details below.
81
+ * `R1`: Difference between observed and modelled for the biggest storm
82
+ * `R1_norm`: Normalized R1 (R1 divided by observed value)
83
+ * `R3`: Average difference between observed and modelled for the three biggest storms
84
+ * `R3_norm`: Normalized R3 (R3 divided by observed value)
85
+ * `error`: Average difference between observed and modelled for all storms
86
+ * `error_norm`: Normalized error (error divided by observed value)
87
+
88
+ ## General metrics
89
+ ### A. Dimensional Statistics:
90
+ #### Mean Error (or Bias)
91
+ $$\langle x_c - x_m \rangle = \langle x_c \rangle - \langle x_m \rangle$$
92
+ #### RMSE (Root Mean Squared Error)
93
+ $$\sqrt{\langle(x_c - x_m)^2\rangle}$$
94
+ #### Mean-Absolute Error (MAE):
95
+ $$\langle |x_c - x_m| \rangle$$
96
+ ### B. Dimentionless Statistics (best closer to 1)
97
+
98
+ #### Performance Scores (PS) or Nash-Sutcliffe Eff (NSE): $$1 - \frac{\langle (x_c - x_m)^2 \rangle}{\langle (x_m - x_R)^2 \rangle}$$
99
+ #### Correlation Coefficient (R):
100
+ $$\frac {\langle x_{m}x_{c}\rangle -\langle x_{m}\rangle \langle x_{c}\rangle }{{\sqrt {\langle x_{m}^{2}\rangle -\langle x_{m}\rangle ^{2}}}{\sqrt {\langle x_{c}^{2}\rangle -\langle x_{c}\rangle ^{2}}}}$$
101
+ #### Kling–Gupta Efficiency (KGE):
102
+ $$1 - \sqrt{(r-1)^2 + b^2 + (g-1)^2}$$
103
+ with :
104
+ * `r` the correlation
105
+ * `b` the modified bias term (see [ref](https://journals.ametsoc.org/view/journals/clim/34/16/JCLI-D-21-0067.1.xml)) $$\frac{\langle x_c \rangle - \langle x_m \rangle}{\sigma_m}$$
106
+ * `g` the std dev term $$\frac{\sigma_c}{\sigma_m}$$
107
+
108
+ #### Lambda index ($\lambda$), values closer to 1 indicate better agreement:
109
+ $$\lambda = 1 - \frac{\sum{(x_c - x_m)^2}}{\sum{(x_m - \overline{x}_m)^2} + \sum{(x_c - \overline{x}_c)^2} + n(\overline{x}_m - \overline{x}_c)^2 + \kappa}$$
110
+ * with `kappa` $$2 \cdot \left| \sum{((x_m - \overline{x}_m) \cdot (x_c - \overline{x}_c))} \right|$$
111
+
112
+ ## Storm metrics
113
+ The functions uses the `match_extremes()` function (detailed below) and returns:
114
+ * `R1`: the error for the biggest storm
115
+ * `R3`: the mean error for the 3 biggest storms
116
+ * `error`: the mean error for all the storms above the threshold.
117
+ * `R1_norm`/`R3_norm`/`error`: Same methodology, but values are in normalised (in %) relatively to the observed peaks.
118
+
119
+
120
+ ### case of NaNs
121
+ The `storm_metrics()` might return:
122
+ ```python
123
+ {'R1': np.nan,
124
+ 'R1_norm': np.nan,
125
+ 'R3': np.nan,
126
+ 'R3_norm': np.nan,
127
+ 'error': np.nan,
128
+ 'error_norm': np.nan}
129
+ ```
130
+ ## Extreme events
131
+
132
+ Example of implementation:
133
+ ```python
134
+ from seastats.storms import match_extremes
135
+ extremes_df = match_extremes(sim, obs, 0.99, cluster = 72)
136
+ extremes_df
137
+ ```
138
+ The modeled peaks are matched with the observed peaks. Function returns a pd.DataFrame of the decreasing observed storm peaks as follows:
139
+
140
+ | time observed | observed | time observed | model | time model | diff | error | error_norm | tdiff |
141
+ |:--------------------|-----------:|:--------------------|---------:|:--------------------|-----------:|----------:|-------------:|--------:|
142
+ | 2022-01-29 19:30:00 | 0.803 | 2022-01-29 19:30:00 | 0.565 | 2022-01-29 17:00:00 | -0.237 | 0.237 | 0.296 | -2.5 |
143
+ | 2022-02-20 20:30:00 | 0.639 | 2022-02-20 20:30:00 | 0.577 | 2022-02-20 20:00:00 | -0.062 | 0.062 | 0.0963 | -0.5 |
144
+ ...
145
+ | 2022-11-27 15:30:00 | 0.386 | 2022-11-27 15:30:00 | 0.400 | 2022-11-27 17:00:00 | 0.014 | 0.014 | 0.036 | 1.5 |
146
+
147
+ with:
148
+ * `diff` the difference between modeled and observed peaks
149
+ * `error` the absolute difference between modeled and observed peaks
150
+ * `tdiff` the time difference between modeled and observed peaks
151
+
152
+ NB: the function uses [pyextremes](https://georgebv.github.io/pyextremes/quickstart/) in the background, with PoT method, using the `quantile` value of the observed signal as physical threshold and passes the `cluster_duration` argument.
153
+
154
+
155
+ this happens when the function `storms/match_extremes.py` couldn't finc concomitent storms for the observed and modeled time series.
156
+
157
+ ## Usage
158
+ see [notebook](/notebooks/example_abed.ipynb) for details
159
+
160
+ get all metrics in a 3 liner:
161
+ ```python
162
+ from seastats import get_stats, GENERAL_METRICS_ALL, STORM_METRICS_ALL
163
+ general = get_stats(sim, obs, metrics = GENERAL_METRICS)
164
+ storm = get_stats(sim, obs, quantile = 0.99, metrics = STORM_METRICS) # we use a different quantile for PoT selection
165
+ pd.DataFrame(dict(general, **storm), index=['abed'])
166
+ ```
167
+
168
+ | | bias | rmse | rms | rms_95 | sim_mean | obs_mean | sim_std | obs_std | nse | lamba | cr | cr_95 | slope | intercept | slope_pp | intercept_pp | mad | madp | madc | kge | R1 | R1_norm | R3 | R3_norm | error | error_norm |
169
+ |:-----|-------:|-------:|------:|---------:|-----------:|-----------:|----------:|----------:|------:|--------:|------:|--------:|--------:|------------:|-----------:|---------------:|------:|-------:|-------:|------:|---------:|----------:|---------:|----------:|----------:|-------------:|
170
+ | abed | -0.007 | 0.086 | 0.086 | 0.088 | -0 | 0.007 | 0.142 | 0.144 | 0.677 | 0.929 | 0.817 | 0.542 | 0.718 | -0.005 | 1.401 | -0.028 | 0.052 | 0.213 | 0.265 | 0.81 | 0.237364 | 0.295719 | 0.147163 | 0.207019 | 0.0938142 | 0.177533 |
171
+
@@ -0,0 +1,142 @@
1
+ # SeaStats
2
+
3
+ `seastats` is a simple package to compare and analyse 2 time series. We use the following convention in this repo:
4
+ * `sim`: modelled surge time series
5
+ * `mod`: observed surge time series
6
+
7
+ The main function is:
8
+
9
+ ```python
10
+ def get_stats(
11
+ sim: Series,
12
+ obs: Series,
13
+ metrics: Sequence[str] = SUGGESTED_METRICS,
14
+ quantile: float = 0,
15
+ cluster: int = 72,
16
+ round: int = -1
17
+ ) -> dict[str, float]
18
+ ```
19
+ Calculates various statistical metrics between the simulated and observed time series data.
20
+ ## Parameters:
21
+ * **sim** (pd.Series). The simulated time series data.
22
+ * **obs** (pd.Series). The observed time series data.
23
+ * **metrics** (list[str]). (Optional) The list of statistical metrics to calculate. If metrics = ["all"], all items in `SUPPORTED_METRICS` will be calculated. Default is all items in `SUGGESTED_METRICS`.
24
+ * **quantile** (float). (Optional) Quantile used to calculate the metrics. Default is `0` (no selection)
25
+ * **cluster** (int). (Optional) Cluster duration for grouping storm events. Default is `72` hours.
26
+ * **round** (int). (Optional) Apply rounding to the results to. Default is no rounding (value is `-1`)
27
+
28
+ Returns a dictionary containing the calculated metrics and their corresponding values. With 2 types of metrics:
29
+ * [The "general" metrics](#general-metrics): All the basic metrics needed for signal comparison (RMSE, RMS, Correlation etc..). See details below
30
+ * `bias`: Bias
31
+ * `rmse`: Root Mean Square Error
32
+ * `rms`: Root Mean Square
33
+ * `rms_95`: Root Mean Square for data points above 95th percentile
34
+ * `sim_mean`: Mean of simulated values
35
+ * `obs_mean`: Mean of observed values
36
+ * `sim_std`: Standard deviation of simulated values
37
+ * `obs_std`: Standard deviation of observed values
38
+ * `mae`: Mean Absolute Error
39
+ * `mse`: Mean Square Error
40
+ * `nse`: Nash-Sutcliffe Efficiency
41
+ * `lamba`: Lambda index
42
+ * `cr`: Pearson Correlation coefficient
43
+ * `cr_95`: Pearson Correlation coefficient for data points above 95th percentile
44
+ * `slope`: Slope of Model/Obs correlation
45
+ * `intercept`: Intercept of Model/Obs correlation
46
+ * `slope_pp`: Slope of Model/Obs correlation of percentiles
47
+ * `intercept_pp`: Intercept of Model/Obs correlation of percentiles
48
+ * `mad`: Mean Absolute Deviation
49
+ * `madp`: Mean Absolute Deviation of percentiles
50
+ * `madc`: `mad + madp`
51
+ * `kge`: Kling–Gupta Efficiency
52
+ * [The storm metrics](#storm-metrics): a PoT selection is done on the observed signal (using the `match_extremes()` function). Function returns the decreasing extreme event peak values for observed and modeled signals (and time lag between events). See details below.
53
+ * `R1`: Difference between observed and modelled for the biggest storm
54
+ * `R1_norm`: Normalized R1 (R1 divided by observed value)
55
+ * `R3`: Average difference between observed and modelled for the three biggest storms
56
+ * `R3_norm`: Normalized R3 (R3 divided by observed value)
57
+ * `error`: Average difference between observed and modelled for all storms
58
+ * `error_norm`: Normalized error (error divided by observed value)
59
+
60
+ ## General metrics
61
+ ### A. Dimensional Statistics:
62
+ #### Mean Error (or Bias)
63
+ $$\langle x_c - x_m \rangle = \langle x_c \rangle - \langle x_m \rangle$$
64
+ #### RMSE (Root Mean Squared Error)
65
+ $$\sqrt{\langle(x_c - x_m)^2\rangle}$$
66
+ #### Mean-Absolute Error (MAE):
67
+ $$\langle |x_c - x_m| \rangle$$
68
+ ### B. Dimentionless Statistics (best closer to 1)
69
+
70
+ #### Performance Scores (PS) or Nash-Sutcliffe Eff (NSE): $$1 - \frac{\langle (x_c - x_m)^2 \rangle}{\langle (x_m - x_R)^2 \rangle}$$
71
+ #### Correlation Coefficient (R):
72
+ $$\frac {\langle x_{m}x_{c}\rangle -\langle x_{m}\rangle \langle x_{c}\rangle }{{\sqrt {\langle x_{m}^{2}\rangle -\langle x_{m}\rangle ^{2}}}{\sqrt {\langle x_{c}^{2}\rangle -\langle x_{c}\rangle ^{2}}}}$$
73
+ #### Kling–Gupta Efficiency (KGE):
74
+ $$1 - \sqrt{(r-1)^2 + b^2 + (g-1)^2}$$
75
+ with :
76
+ * `r` the correlation
77
+ * `b` the modified bias term (see [ref](https://journals.ametsoc.org/view/journals/clim/34/16/JCLI-D-21-0067.1.xml)) $$\frac{\langle x_c \rangle - \langle x_m \rangle}{\sigma_m}$$
78
+ * `g` the std dev term $$\frac{\sigma_c}{\sigma_m}$$
79
+
80
+ #### Lambda index ($\lambda$), values closer to 1 indicate better agreement:
81
+ $$\lambda = 1 - \frac{\sum{(x_c - x_m)^2}}{\sum{(x_m - \overline{x}_m)^2} + \sum{(x_c - \overline{x}_c)^2} + n(\overline{x}_m - \overline{x}_c)^2 + \kappa}$$
82
+ * with `kappa` $$2 \cdot \left| \sum{((x_m - \overline{x}_m) \cdot (x_c - \overline{x}_c))} \right|$$
83
+
84
+ ## Storm metrics
85
+ The functions uses the `match_extremes()` function (detailed below) and returns:
86
+ * `R1`: the error for the biggest storm
87
+ * `R3`: the mean error for the 3 biggest storms
88
+ * `error`: the mean error for all the storms above the threshold.
89
+ * `R1_norm`/`R3_norm`/`error`: Same methodology, but values are in normalised (in %) relatively to the observed peaks.
90
+
91
+
92
+ ### case of NaNs
93
+ The `storm_metrics()` might return:
94
+ ```python
95
+ {'R1': np.nan,
96
+ 'R1_norm': np.nan,
97
+ 'R3': np.nan,
98
+ 'R3_norm': np.nan,
99
+ 'error': np.nan,
100
+ 'error_norm': np.nan}
101
+ ```
102
+ ## Extreme events
103
+
104
+ Example of implementation:
105
+ ```python
106
+ from seastats.storms import match_extremes
107
+ extremes_df = match_extremes(sim, obs, 0.99, cluster = 72)
108
+ extremes_df
109
+ ```
110
+ The modeled peaks are matched with the observed peaks. Function returns a pd.DataFrame of the decreasing observed storm peaks as follows:
111
+
112
+ | time observed | observed | time observed | model | time model | diff | error | error_norm | tdiff |
113
+ |:--------------------|-----------:|:--------------------|---------:|:--------------------|-----------:|----------:|-------------:|--------:|
114
+ | 2022-01-29 19:30:00 | 0.803 | 2022-01-29 19:30:00 | 0.565 | 2022-01-29 17:00:00 | -0.237 | 0.237 | 0.296 | -2.5 |
115
+ | 2022-02-20 20:30:00 | 0.639 | 2022-02-20 20:30:00 | 0.577 | 2022-02-20 20:00:00 | -0.062 | 0.062 | 0.0963 | -0.5 |
116
+ ...
117
+ | 2022-11-27 15:30:00 | 0.386 | 2022-11-27 15:30:00 | 0.400 | 2022-11-27 17:00:00 | 0.014 | 0.014 | 0.036 | 1.5 |
118
+
119
+ with:
120
+ * `diff` the difference between modeled and observed peaks
121
+ * `error` the absolute difference between modeled and observed peaks
122
+ * `tdiff` the time difference between modeled and observed peaks
123
+
124
+ NB: the function uses [pyextremes](https://georgebv.github.io/pyextremes/quickstart/) in the background, with PoT method, using the `quantile` value of the observed signal as physical threshold and passes the `cluster_duration` argument.
125
+
126
+
127
+ this happens when the function `storms/match_extremes.py` couldn't finc concomitent storms for the observed and modeled time series.
128
+
129
+ ## Usage
130
+ see [notebook](/notebooks/example_abed.ipynb) for details
131
+
132
+ get all metrics in a 3 liner:
133
+ ```python
134
+ from seastats import get_stats, GENERAL_METRICS_ALL, STORM_METRICS_ALL
135
+ general = get_stats(sim, obs, metrics = GENERAL_METRICS)
136
+ storm = get_stats(sim, obs, quantile = 0.99, metrics = STORM_METRICS) # we use a different quantile for PoT selection
137
+ pd.DataFrame(dict(general, **storm), index=['abed'])
138
+ ```
139
+
140
+ | | bias | rmse | rms | rms_95 | sim_mean | obs_mean | sim_std | obs_std | nse | lamba | cr | cr_95 | slope | intercept | slope_pp | intercept_pp | mad | madp | madc | kge | R1 | R1_norm | R3 | R3_norm | error | error_norm |
141
+ |:-----|-------:|-------:|------:|---------:|-----------:|-----------:|----------:|----------:|------:|--------:|------:|--------:|--------:|------------:|-----------:|---------------:|------:|-------:|-------:|------:|---------:|----------:|---------:|----------:|----------:|-------------:|
142
+ | abed | -0.007 | 0.086 | 0.086 | 0.088 | -0 | 0.007 | 0.142 | 0.144 | 0.677 | 0.929 | 0.817 | 0.542 | 0.718 | -0.005 | 1.401 | -0.028 | 0.052 | 0.213 | 0.265 | 0.81 | 0.237364 | 0.295719 | 0.147163 | 0.207019 | 0.0938142 | 0.177533 |
@@ -0,0 +1,123 @@
1
+ [tool.poetry]
2
+ name = "seastats"
3
+ version = "0.1.0"
4
+ description = "package for metocean statistics"
5
+ authors = [
6
+ "tomsail <saillour.thomas@gmail.com>",
7
+ "pmav99 <pmav99@gmail.com>",
8
+ ]
9
+ license = 'EUPL-1.2'
10
+ readme = "README.md"
11
+ repository = "https://github.com/oceanmodeling/seastats.git"
12
+ classifiers = [
13
+ "Operating System :: OS Independent",
14
+ "Development Status :: 4 - Beta",
15
+ "Environment :: Other Environment",
16
+ "Intended Audience :: Developers",
17
+ "Intended Audience :: Science/Research",
18
+ "Topic :: Scientific/Engineering :: Physics",
19
+ "Programming Language :: Python",
20
+ "Programming Language :: Python :: 3",
21
+ "Programming Language :: Python :: 3.10",
22
+ "Programming Language :: Python :: 3.11",
23
+ "Programming Language :: Python :: 3.12",
24
+ "Programming Language :: Python :: 3.13",
25
+ ]
26
+
27
+ [tool.poetry.dependencies]
28
+ python = "^3.10"
29
+ numpy = "*"
30
+ pandas = {version = "*", extras = ["pyarrow", "performance"]}
31
+ pyextremes = "*"
32
+
33
+ [tool.poetry.group.dev.dependencies]
34
+ covdefaults = "*"
35
+ hvplot = "*"
36
+ ipykernel = "*"
37
+ mypy = "*"
38
+ pytest = "*"
39
+ pytest-cov = "*"
40
+ pandas-stubs = "*"
41
+
42
+ [build-system]
43
+ requires = ["poetry-core"]
44
+ build-backend = "poetry.core.masonry.api"
45
+
46
+ [tool.mypy]
47
+ python_version = "3.10"
48
+ plugins = [
49
+ "numpy.typing.mypy_plugin"
50
+ ]
51
+ # ignore_missing_imports = true
52
+ show_column_numbers = true
53
+ show_error_codes = true
54
+ show_error_context = true
55
+ warn_no_return = true
56
+ warn_redundant_casts = true
57
+ warn_return_any = true
58
+ warn_unreachable = true
59
+ warn_unused_ignores = true
60
+ strict = true
61
+ disable_error_code = [ ]
62
+ enable_error_code = [
63
+ "comparison-overlap",
64
+ "explicit-override",
65
+ "ignore-without-code",
66
+ "no-any-return",
67
+ "no-any-unimported",
68
+ "no-untyped-call",
69
+ "no-untyped-def",
70
+ "possibly-undefined",
71
+ "redundant-cast",
72
+ "redundant-expr",
73
+ "redundant-self",
74
+ "truthy-bool",
75
+ "truthy-iterable",
76
+ "type-arg",
77
+ "unimported-reveal",
78
+ "unreachable",
79
+ "unused-ignore",
80
+ ]
81
+
82
+ # mypy per-module options:
83
+ [[tool.mypy.overrides]]
84
+ module = "tests.*"
85
+ disallow_untyped_defs = true
86
+
87
+ [tool.ruff]
88
+ target-version = "py310"
89
+ line-length = 108
90
+ lint.select = [
91
+ "E", # pycodestyle
92
+ "F", # pyflakes
93
+ "C90", # mccabe
94
+ "A", # flake8-builtins
95
+ "COM", # flake8-commas
96
+ # "UP", # pyupgrade
97
+ # "YTT", # flake-2020
98
+ # "S", # floke8-bandit
99
+ # "BLE", # flake8-blind-except
100
+ # "B", # flake8-bugbear
101
+ # "T20", # flake8-print
102
+ # "PD", # pandas-vet
103
+ # "NPY", # numpy-specific rules
104
+ # "RUF", # ruff-specific rules
105
+ # "D", # pydocstyle
106
+ # "I", # isort
107
+ # "N", # pep8-naming
108
+ ]
109
+ lint.ignore = [
110
+ "E501", # line-too-long
111
+ "D103", # undocumented-public-function
112
+ "PD901", # pandas-df-variable-name
113
+ ]
114
+
115
+ [tool.coverage.run]
116
+ plugins = ["covdefaults"]
117
+ source = ["seastats"]
118
+ omit = []
119
+ parallel = true
120
+ sigterm = true
121
+
122
+ [tool.coverage.report]
123
+ fail_under = 86.8
@@ -0,0 +1,188 @@
1
+ from __future__ import annotations
2
+
3
+ import logging
4
+ from collections.abc import Sequence
5
+
6
+ import numpy as np
7
+ import pandas as pd
8
+
9
+ from seastats.stats import get_bias
10
+ from seastats.stats import get_corr
11
+ from seastats.stats import get_kge
12
+ from seastats.stats import get_lambda
13
+ from seastats.stats import get_mad
14
+ from seastats.stats import get_madc
15
+ from seastats.stats import get_madp
16
+ from seastats.stats import get_mae
17
+ from seastats.stats import get_mse
18
+ from seastats.stats import get_nse
19
+ from seastats.stats import get_rms
20
+ from seastats.stats import get_rmse
21
+ from seastats.stats import get_slope_intercept
22
+ from seastats.stats import get_slope_intercept_pp
23
+ from seastats.storms import match_extremes
24
+
25
+ logger = logging.getLogger(__name__)
26
+
27
+ GENERAL_METRICS_ALL = [
28
+ "bias",
29
+ "rmse",
30
+ "mae",
31
+ "mse",
32
+ "rms",
33
+ "sim_mean",
34
+ "obs_mean",
35
+ "sim_std",
36
+ "obs_std",
37
+ "nse",
38
+ "lamba",
39
+ "cr",
40
+ "slope",
41
+ "intercept",
42
+ "slope_pp",
43
+ "intercept_pp",
44
+ "mad",
45
+ "madp",
46
+ "madc",
47
+ "kge",
48
+ ]
49
+ GENERAL_METRICS = ["bias", "rms", "rmse", "cr", "nse", "kge"]
50
+ STORM_METRICS = ["R1", "R3", "error"]
51
+ STORM_METRICS_ALL = ["R1", "R1_norm", "R3", "R3_norm", "error", "error_norm"]
52
+
53
+ SUGGESTED_METRICS = sorted(GENERAL_METRICS + STORM_METRICS)
54
+ SUPPORTED_METRICS = sorted(GENERAL_METRICS_ALL + STORM_METRICS_ALL)
55
+
56
+
57
+ def get_stats(
58
+ sim: pd.Series[float],
59
+ obs: pd.Series[float],
60
+ metrics: Sequence[str] = SUGGESTED_METRICS,
61
+ quantile: float = 0,
62
+ cluster: int = 72,
63
+ round: int = -1, # noqa: A002
64
+ ) -> dict[str, float]:
65
+ """
66
+ Calculates various statistical metrics between the simulated and observed time series data.
67
+
68
+ :param pd.Series sim: The simulated time series data.
69
+ :param pd.Series obs: The observed time series data.
70
+ :param list[str] metrics: (Optional) The list of statistical metrics to calculate. If metrics = ["all"], all items in SUPPORTED_METRICS will be calculated. Default is all items in SUGGESTED_METRICS.
71
+ :param float quantile: (Optional) Quantile used to calculate the metrics. Default is 0 (no selection)
72
+ :param int cluster: (Optional) Cluster duration for grouping storm events. Default is 72 hours.
73
+ :param int round: (Optional) Apply rounding to the results to. Default is no rounding (value is -1)
74
+
75
+ :return dict stats: dictionary containing the calculated metrics.
76
+
77
+ The dictionary contains the following keys and their corresponding values:
78
+
79
+ - `bias`: The bias between the simulated and observed time series data.
80
+ - `rmse`: The Root Mean Square Error between the simulated and observed time series data.
81
+ - `mae`: The Mean Absolute Error the simulated and observed time series data.
82
+ - `mse`: The Mean Square Error the simulated and observed time series data.
83
+ - `rms`: The raw mean square error between the simulated and observed time series data.
84
+ - `sim_mean`: The mean of the simulated time series data.
85
+ - `obs_mean`: The mean of the observed time series data.
86
+ - `sim_std`: The standard deviation of the simulated time series data.
87
+ - `obs_std`: The standard deviation of the observed time series data.
88
+ - `nse`: The Nash-Sutcliffe efficiency between the simulated and observed time series data.
89
+ - `lamba`: The lambda statistic between the simulated and observed time series data.
90
+ - `cr`: The correlation coefficient between the simulated and observed time series data.
91
+ - `slope`: The slope of the linear regression between the simulated and observed time series data.
92
+ - `intercept`: The intercept of the linear regression between the simulated and observed time series data.
93
+ - `slope_pp`: The slope of the linear regression between the percentiles of the simulated and observed time series data.
94
+ - `intercept_pp`: The intercept of the linear regression between the percentiles of the simulated and observed time series data.
95
+ - `mad`: The median absolute deviation of the simulated time series data from its median.
96
+ - `madp`: The median absolute deviation of the simulated time series data from its median, calculated using the percentiles of the observed time series data.
97
+ - `madc`: The median absolute deviation of the simulated time series data from its median, calculated by adding `mad` to `madp`
98
+ - `kge`: The Kling-Gupta efficiency between the simulated and observed time series data.
99
+ - `R1`: Difference between observed and modelled for the biggest storm
100
+ - `R1_norm`: Normalized R1 (R1 divided by observed value)
101
+ - `R3`: Average difference between observed and modelled for the three biggest storms
102
+ - `R3_norm`: Normalized R3 (R3 divided by observed value)
103
+ - `error`: Average difference between observed and modelled for all storms
104
+ - `error_norm`: Normalized error (error divided by observed value)
105
+ """
106
+ if not isinstance(metrics, list):
107
+ raise ValueError("metrics must be a list")
108
+
109
+ if metrics == ["all"]:
110
+ metrics = SUPPORTED_METRICS
111
+
112
+ if not np.any([m in SUPPORTED_METRICS for m in metrics]):
113
+ raise ValueError("metrics must be a list of supported variables in SUPPORTED_METRICS or ['all']")
114
+
115
+ # Storm metrics part with PoT Selection
116
+ if np.any([m in STORM_METRICS_ALL for m in metrics]):
117
+ extreme_df = match_extremes(sim, obs, quantile=quantile, cluster=cluster)
118
+ else:
119
+ extreme_df = pd.DataFrame() # Just to make mypy happy
120
+
121
+ if quantile: # signal subsetting is only to be done on general metrics
122
+ sim = sim[sim > sim.quantile(quantile)]
123
+ obs = obs[obs > obs.quantile(quantile)]
124
+
125
+ stats = {}
126
+
127
+ for metric in metrics:
128
+ match metric:
129
+ case "bias":
130
+ stats["bias"] = get_bias(sim, obs)
131
+ case "rmse":
132
+ stats["rmse"] = get_rmse(sim, obs)
133
+ case "rms":
134
+ stats["rms"] = get_rms(sim, obs)
135
+ case "sim_mean":
136
+ stats["sim_mean"] = sim.mean()
137
+ case "obs_mean":
138
+ stats["obs_mean"] = obs.mean()
139
+ case "sim_std":
140
+ stats["sim_std"] = sim.std()
141
+ case "obs_std":
142
+ stats["obs_std"] = obs.std()
143
+ case "mae":
144
+ stats["mae"] = get_mae(sim, obs)
145
+ case "mse":
146
+ stats["mse"] = get_mse(sim, obs)
147
+ case "nse":
148
+ stats["nse"] = get_nse(sim, obs)
149
+ case "lamba":
150
+ stats["lamba"] = get_lambda(sim, obs)
151
+ case "cr":
152
+ stats["cr"] = get_corr(sim, obs)
153
+ case "slope":
154
+ stats["slope"], _ = get_slope_intercept(sim, obs)
155
+ case "intercept":
156
+ _, stats["intercept"] = get_slope_intercept(sim, obs)
157
+ case "slope_pp":
158
+ stats["slope_pp"], _ = get_slope_intercept_pp(sim, obs)
159
+ case "intercept_pp":
160
+ _, stats["intercept_pp"] = get_slope_intercept_pp(sim, obs)
161
+ case "mad":
162
+ stats["mad"] = get_mad(sim, obs)
163
+ case "madp":
164
+ stats["madp"] = get_madp(sim, obs)
165
+ case "madc":
166
+ stats["madc"] = get_madc(sim, obs)
167
+ case "kge":
168
+ stats["kge"] = get_kge(sim, obs)
169
+ case "R1":
170
+ stats["R1"] = extreme_df["error"].iloc[0]
171
+ case "R1_norm":
172
+ stats["R1_norm"] = extreme_df["error_norm"].iloc[0]
173
+ case "R3":
174
+ stats["R3"] = extreme_df["error"].iloc[0:3].mean()
175
+ case "R3_norm":
176
+ stats["R3_norm"] = extreme_df["error_norm"].iloc[0:3].mean()
177
+ case "error":
178
+ stats["error"] = extreme_df["error"].mean()
179
+ case "error_norm":
180
+ stats["error_norm"] = extreme_df["error_norm"].mean()
181
+ else:
182
+ logger.info("no storm metric specified")
183
+
184
+ if round > 0:
185
+ for metric in metrics:
186
+ stats[metric] = np.round(stats[metric], round)
187
+
188
+ return stats
File without changes
@@ -0,0 +1,149 @@
1
+ from __future__ import annotations
2
+
3
+ import logging
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+
8
+ logger = logging.getLogger(__name__)
9
+
10
+
11
+ def get_bias(sim: pd.Series[float], obs: pd.Series[float]) -> float:
12
+ return float(sim.mean() - obs.mean())
13
+
14
+
15
+ def get_mse(sim: pd.Series[float], obs: pd.Series[float]) -> float:
16
+ return float(np.square(np.subtract(obs, sim)).mean())
17
+
18
+
19
+ def get_rmse(sim: pd.Series[float], obs: pd.Series[float]) -> float:
20
+ return float(np.sqrt(get_mse(sim, obs)))
21
+
22
+
23
+ def get_mae(sim: pd.Series[float], obs: pd.Series[float]) -> float:
24
+ return float(np.abs(np.subtract(obs, sim)).mean())
25
+
26
+
27
+ def get_mad(sim: pd.Series[float], obs: pd.Series[float]) -> float:
28
+ return float(np.abs(np.subtract(obs, sim)).std())
29
+
30
+
31
+ def get_madp(sim: pd.Series[float], obs: pd.Series[float]) -> float:
32
+ pc1, pc2 = get_percentiles(sim, obs)
33
+ return get_mad(pc1, pc2)
34
+
35
+
36
+ def get_madc(sim: pd.Series[float], obs: pd.Series[float]) -> float:
37
+ madp = get_madp(sim, obs)
38
+ return get_mad(sim, obs) + madp
39
+
40
+
41
+ def get_rms(sim: pd.Series[float], obs: pd.Series[float]) -> float:
42
+ crmsd = ((sim - sim.mean()) - (obs - obs.mean())) ** 2
43
+ return float(np.sqrt(crmsd.mean()))
44
+
45
+
46
+ def get_corr(sim: pd.Series[float], obs: pd.Series[float]) -> float:
47
+ return float(sim.corr(obs))
48
+
49
+
50
+ def get_nse(sim: pd.Series[float], obs: pd.Series[float]) -> float:
51
+ nse = 1 - np.nansum(np.subtract(obs, sim) ** 2) / np.nansum((obs - float(np.nanmean(obs))) ** 2)
52
+ return float(nse)
53
+
54
+
55
+ def get_lambda(sim: pd.Series[float], obs: pd.Series[float]) -> float:
56
+ Xmean = float(np.nanmean(obs))
57
+ Ymean = float(np.nanmean(sim))
58
+ nObs = len(obs)
59
+ corr = get_corr(sim, obs)
60
+ if corr >= 0:
61
+ kappa = 0
62
+ else:
63
+ kappa = 2 * abs(np.nansum((obs - Xmean) * (sim - Ymean)))
64
+
65
+ numerator = np.nansum((obs - sim) ** 2)
66
+ denominator = (
67
+ np.nansum((obs - Xmean) ** 2)
68
+ + np.nansum((sim - Ymean) ** 2)
69
+ + nObs * ((Xmean - Ymean) ** 2)
70
+ + kappa
71
+ )
72
+ lambda_index = 1 - numerator / denominator
73
+ return float(lambda_index)
74
+
75
+
76
+ def get_kge(sim: pd.Series[float], obs: pd.Series[float]) -> float:
77
+ corr = get_corr(sim, obs)
78
+ b = (sim.mean() - obs.mean()) / obs.std()
79
+ g = sim.std() / obs.std()
80
+ return float(1 - np.sqrt((corr - 1) ** 2 + b**2 + (g - 1) ** 2))
81
+
82
+
83
+ def truncate_seconds(ts: pd.Series[float]) -> pd.Series[float]:
84
+ df = pd.DataFrame({"time": ts.index, "value": ts.values})
85
+ df = df.assign(time=df.time.dt.floor("min"))
86
+ if df.time.duplicated().any():
87
+ # There are duplicates. Keep the first datapoint per minute.
88
+ msg = "Duplicate timestamps have been detected after the truncation of seconds. Keeping the first datapoint per minute"
89
+ logger.warning(msg)
90
+ df = df.iloc[df.time.drop_duplicates().index].reset_index(drop=True)
91
+ df.index = df.time # type: ignore[assignment]
92
+ df = df.drop("time", axis=1)
93
+ ts = pd.Series(index=df.index, data=df.value)
94
+ return ts
95
+
96
+
97
+ def align_ts(
98
+ sim: pd.Series[float],
99
+ obs: pd.Series[float],
100
+ resample_to_model: bool = True,
101
+ ) -> tuple[pd.Series[float], pd.Series[float]]:
102
+ # requisite: obs is the observation time series
103
+ obs = obs.dropna()
104
+ obs = truncate_seconds(obs)
105
+ if resample_to_model:
106
+ freq = sim.index.to_series().diff().median()
107
+ obs = obs.resample(pd.Timedelta(freq)).mean()
108
+ sim_, obs_ = sim.align(obs, axis=0)
109
+ nan_mask1 = pd.isna(sim_)
110
+ nan_mask2 = pd.isna(obs_)
111
+ nan_mask = np.logical_or(nan_mask1, nan_mask2)
112
+ sim_ = sim_[~nan_mask]
113
+ obs_ = obs_[~nan_mask]
114
+ return sim_, obs_
115
+
116
+
117
+ def get_percentiles(
118
+ sim: pd.Series[float],
119
+ obs: pd.Series[float],
120
+ higher_tail: bool = False,
121
+ ) -> tuple[pd.Series[float], pd.Series[float]]:
122
+ x = np.arange(0, 0.99, 0.01)
123
+ if higher_tail:
124
+ x = np.hstack([x, np.arange(0.99, 1, 0.001)])
125
+
126
+ pc_sim = sim.quantile(x).to_numpy()
127
+ pc_obs = obs.quantile(x).to_numpy()
128
+
129
+ return pd.Series(pc_sim), pd.Series(pc_obs)
130
+
131
+
132
+ def get_slope_intercept(sim: pd.Series[float], obs: pd.Series[float]) -> tuple[float, float]:
133
+ # Calculate means of x and y
134
+ x_mean = float(np.mean(obs))
135
+ y_mean = float(np.mean(sim))
136
+
137
+ numerator = np.sum((obs - x_mean) * (sim - y_mean))
138
+ denominator = np.sum((obs - x_mean) ** 2)
139
+
140
+ # Calculate slope (A) and intercept (B) in A*X + B
141
+ slope = numerator / denominator
142
+ intercept = y_mean - slope * x_mean
143
+ return slope, intercept
144
+
145
+
146
+ def get_slope_intercept_pp(sim: pd.Series[float], obs: pd.Series[float]) -> tuple[float, float]:
147
+ pc1, pc2 = get_percentiles(sim, obs)
148
+ slope, intercept = get_slope_intercept(pc1, pc2)
149
+ return slope, intercept
@@ -0,0 +1,70 @@
1
+ from __future__ import annotations
2
+
3
+ import typing as T
4
+
5
+ import numpy as np
6
+ import pandas as pd
7
+ from pyextremes import get_extremes
8
+
9
+
10
+ def match_extremes(
11
+ sim: pd.Series[float],
12
+ obs: pd.Series[float],
13
+ quantile: float,
14
+ cluster: int = 72,
15
+ ) -> pd.DataFrame:
16
+ """
17
+ Calculate metrics for comparing simulated and observed storm events.
18
+ Parameters:
19
+ - sim (pd.Series): Simulated storm series.
20
+ - obs (pd.Series): Observed storm series.
21
+ - quantile (float): Quantile value for defining extreme events.
22
+ - cluster (int, optional): Cluster duration for grouping storm events. Default is 72 hours.
23
+
24
+ Returns:
25
+ - df: pd.DataFrame of the matched extremes between the observed and modeled pd.Series.
26
+ with following columns:
27
+ * `observed`: observed extreme event value
28
+ * `time observed`: observed extreme event time
29
+ * `model`: modeled extreme event value
30
+ * `time model`: modeled extreme event time
31
+ * `diff`: difference between model and observed
32
+ * `error`: absolute difference between model and observed
33
+ * `error_norm`: normalised difference between model and observed
34
+ * `tdiff`: time difference between model and observed (in hours)
35
+
36
+ !Important: The modeled values are matched on the observed events calculated by POT analysis.
37
+ The user needs to be mindful about the order of the observed and modeled pd.Series.
38
+ """
39
+ # resample observation to 1H time series
40
+ obs = obs.resample("1h").mean().shift(freq="30min")
41
+ # get observed extremes
42
+ ext = get_extremes(obs, "POT", threshold=obs.quantile(quantile), r=f"{cluster}h")
43
+ ext_values_dict: dict[str, T.Any] = {}
44
+ ext_values_dict["observed"] = ext.values
45
+ ext_values_dict["time observed"] = ext.index.values
46
+ #
47
+ max_in_window = []
48
+ tmax_in_window = []
49
+ # match simulated values with observed events
50
+ for it, itime in enumerate(ext.index):
51
+ snippet = sim[itime - pd.Timedelta(hours=cluster / 2) : itime + pd.Timedelta(hours=cluster / 2)]
52
+ try:
53
+ tmax_in_window.append(snippet.index[int(snippet.argmax())])
54
+ max_in_window.append(snippet.max())
55
+ except Exception:
56
+ tmax_in_window.append(itime)
57
+ max_in_window.append(np.nan)
58
+ ext_values_dict["model"] = max_in_window
59
+ ext_values_dict["time model"] = tmax_in_window
60
+ #
61
+ df = pd.DataFrame(ext_values_dict)
62
+ df = df.dropna(subset="model")
63
+ df = df.sort_values("observed", ascending=False)
64
+ df["diff"] = df["model"] - df["observed"]
65
+ df["error"] = abs(df["diff"])
66
+ df["error_norm"] = abs(df["diff"] / df["observed"])
67
+ df["tdiff"] = df["time model"] - df["time observed"]
68
+ df["tdiff"] = df["tdiff"].apply(lambda x: x.total_seconds() / 3600)
69
+ df = df.set_index("time observed")
70
+ return df