UFCData 0.7.0__tar.gz → 0.7.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: UFCData
3
- Version: 0.7.0
3
+ Version: 0.7.2
4
4
  Summary: Obtain UFC data and functions to manipulate it
5
5
  Requires-Python: >=3.10
6
6
  License-File: LICENSE.md
@@ -1,11 +1,14 @@
1
1
  # UFCData
2
2
 
3
+
3
4
  An open-source Python package for accessing, manipulating, and analyzing UFC data.
4
5
 
5
6
  https://pypi.org/project/UFCData/
6
7
 
7
8
  ## Table of Contents
8
9
 
10
+
11
+
9
12
  - [Installation](#installation)
10
13
  - [Updating the Package](#updating-the-package)
11
14
 
@@ -30,7 +33,7 @@ Model Evaluation
30
33
  - [Train Test Split](#train-test-split)
31
34
  - [Get Expanding Window](#get-expanding-window)
32
35
  - [Plotting Accuracy](#plotting-accuracy)
33
- - [Expanding Window Baseline Example](#expanding-window-cv-and-final-test-example)
36
+ - [Expanding Window Baseline Example](#expanding-window-baseline-example)
34
37
 
35
38
 
36
39
  Data Sources
@@ -38,6 +41,12 @@ Data Sources
38
41
  - [Discrepancies in Data](#discrepancies-in-data)
39
42
  - [Update Frequency](#update-frequency)
40
43
 
44
+ <br>
45
+
46
+ ---
47
+
48
+ <br>
49
+
41
50
  ## Installation
42
51
 
43
52
  ```bash
@@ -50,18 +59,23 @@ pip install ufcdata
50
59
  pip install --no-cache-dir --upgrade ufcdata
51
60
  ```
52
61
 
62
+ ---
63
+
64
+ <br>
65
+
53
66
  ## Loading Data
54
67
 
55
68
  ```python
56
69
  import ufcdata as ufc
57
70
  data = ufc.get_data()
58
71
 
59
- ## Obtain fighter_bio dataframe
60
72
  fighter_bio = data["fighter_bio"].copy()
61
73
  ```
62
74
 
63
75
  Note that data in here and other functions below reference data that is available prior to the data.
64
76
 
77
+ <br>
78
+
65
79
  ## Search
66
80
 
67
81
  Due to various event naming conventions and fighters sharing the same name, the primary keys for the dataframes are links, which can be difficult for humans to interpret. The `search()` function uses fuzzy string matching to make it easier to find fighters, events, and other records.
@@ -77,28 +91,19 @@ search(data, column, query, matches=1)
77
91
  * `query` — The name or text you are searching for.
78
92
  * `matches` — The number of closest matches to return. Defaults to `1`.
79
93
 
94
+ ### Returns
95
+ The results are returned as a dataframe, with the closest match appearing first.
96
+
80
97
  ### Example
81
98
 
82
99
  ```python
83
100
  john_jones = ufc.search(fighter_bio, "Name", "Jon Jones", 3)
84
101
  ```
85
102
 
86
- This searches the `"Name"` column of the `fighter_bio` dataframe for the three closest matches to `"Jon Jones"`.
87
103
 
88
- The results are returned as a dataframe, with the closest match appearing first.
89
-
90
- ### Returns
91
-
92
-
93
- ```text
94
- Name Nickname Height Weight Reach Stance Total W Total L Fighter Link
95
- Jon Jones Bones 76.0 248.0 84.0 Orthodox 28 1 ufcstats.com/...
96
- Roshaun Jones NaN 68.0 135.0 NaN NaN 2 6 ufcstats.com/...
97
- Mason Jones The Dragon 70.0 155.0 74.0 Orthodox 18 2 ufcstats.com/...
98
- ```
99
-
100
- This is useful when the exact value stored in the dataset is unknown or when there are multiple similar names. The same function can be used with any dataframe and column containing searchable text.
104
+ ---
101
105
 
106
+ <br>
102
107
 
103
108
  ## Elo
104
109
 
@@ -123,8 +128,8 @@ ufc.get_elo(fight_df, fighter_df, r=1500, k=30, s=400)
123
128
 
124
129
  The function returns a tuple containing:
125
130
 
126
- 1. `fighter_elo` — A dictionary mapping each fighter's UFCStats link to their final Elo rating.
127
- 2. `elo` — The original `fights_df` dataframe with `"Fighter 1 Elo"` and `"Fighter 2 Elo"` columns added. These columns contain each fighter's Elo rating immediately prior to the corresponding fight.
131
+ 1. `current_elo` — A dictionary mapping each fighter's UFCStats link to their final Elo rating.
132
+ 2. `past_elo` — The original `fights_df` dataframe with `"Fighter 1 Elo"` and `"Fighter 2 Elo"` columns added. These columns contain each fighter's Elo rating immediately prior to the corresponding fight.
128
133
 
129
134
 
130
135
  ### Example
@@ -143,13 +148,17 @@ current_elo, past_elo = ufc.get_elo(
143
148
  )
144
149
  ```
145
150
 
146
- The resulting `fight_elo` dataframe can then be used to examine or incorporate pre-fight Elo ratings into analysis and predictive models.
147
-
148
151
 
149
152
  ### Important
150
153
 
151
154
  `fight_df` should be sorted from most recent to least recent before being passed to `get_elo()`. The function reverses the dataframe internally to process fights chronologically and returns the resulting data in the original order.
152
155
 
156
+ <br>
157
+
158
+ ---
159
+
160
+ <br>
161
+
153
162
  ## Fighter History
154
163
 
155
164
  The function `get_fighter_history()` retrieves all UFC fights for a specific fighter and formats the results from the fighter's perspective. The requested fighter is always represented as `"Fighter 1"`, regardless of which side of the original fight dataframe they appeared on.
@@ -198,6 +207,8 @@ jon_jones_history = ufc.get_fighter_history(
198
207
 
199
208
  This function is useful for analyzing an individual fighter's career, constructing fighter-level features, or preparing historical data for predictive modeling.
200
209
 
210
+ <br>
211
+
201
212
  ## Fighter Statistic
202
213
 
203
214
 
@@ -235,11 +246,7 @@ get_fighter_statistic(
235
246
 
236
247
  ### Returns
237
248
 
238
- The function returns two DataFrames:
239
-
240
- ```python
241
- current_statistic, past_statistic = get_fighter_statistic(...)
242
- ```
249
+ The function returns a tuple containing:
243
250
 
244
251
  #### `current_statistic`
245
252
 
@@ -303,24 +310,28 @@ jon_jones_current, jon_jones_past = ufc.get_fighter_statistic(
303
310
  fights_df,
304
311
  rounds_df
305
312
  )
306
-
307
- jon_jones_current
308
313
  ```
314
+
315
+
316
+ ---
317
+
318
+ <br>
319
+
309
320
  ## Odds Convert
310
321
 
311
322
  The function `convert_odds()` converts American betting odds from strings into integers or floating point numbers.
312
323
 
313
324
  ```python
314
- convert_odds(odds, probability)
325
+ convert_odds(odds)
315
326
  ```
316
327
 
317
328
  ### Parameters
318
329
 
319
- * `odds` — American betting odds as a string.
330
+ * `odds` — American betting odds as a string, integer, float or None.
320
331
 
321
332
  ### Returns
322
333
 
323
- If probability=True, it returns a `float` representing the implied probability of the betting odds.
334
+ Returns a `float` representing the implied probability of the betting odds.
324
335
 
325
336
  For positive American odds:
326
337
 
@@ -333,7 +344,6 @@ For negative American odds:
333
344
  ```text
334
345
  Probability = -odds / (-odds + 100)
335
346
  ```
336
- If probability=False, takes the string and converts it to an integer.
337
347
 
338
348
  ### Example
339
349
 
@@ -344,10 +354,13 @@ data = ufc.get_data()
344
354
 
345
355
  fights_df = data["past_fights"].copy()
346
356
 
347
- fights_df["Fighter 1 Odds"] = fights_df["Fighter 1 Odds"].apply(ufc.convert_odds)
348
- fights_df["Fighter 1 Probability"] = fights_df["Fighter 1 Odds"].apply(ufc.convert_odds, probability=True)
349
- fights_df
357
+ fights_df["Fighter 1 Probability"] = fights_df["Fighter 1 Odds"].apply(
358
+ ufc.convert_odds,
359
+ )
350
360
  ```
361
+
362
+ <br>
363
+
351
364
  ## Gender Convert
352
365
 
353
366
  The function `convert_gender()` converts gender labels into binary numeric values, with `"Male"` represented as `1` and `"Female"` represented as `0`.
@@ -375,9 +388,10 @@ data = ufc.get_data()
375
388
 
376
389
  fights_df = data["past_fights"].copy()
377
390
  fights_df["Gender"] = fights_df["Gender"].apply(ufc.convert_gender)
378
- fights_df
379
391
  ```
380
392
 
393
+ <br>
394
+
381
395
  ## Weight Convert
382
396
 
383
397
  The function `convert_weight()` converts UFC weight class labels into their corresponding weight limits in pounds.
@@ -418,10 +432,10 @@ data = ufc.get_data()
418
432
 
419
433
  past_fights = data["past_fights"].copy()
420
434
  past_fights["Weight Class"] = past_fights["Weight Class"].apply(ufc.convert_weight)
421
-
422
- past_fights
423
435
  ```
424
436
 
437
+ <br>
438
+
425
439
  ## Mirror
426
440
 
427
441
  The function `mirror()` creates mirrored versions of a DataFrame by swapping corresponding pairs of columns. This is useful when analyzing UFC fights from both fighters' perspectives.
@@ -453,7 +467,16 @@ import ufcdata as ufc
453
467
  data = ufc.get_data()
454
468
 
455
469
  fights_df = data["past_fights"].copy()
456
- fights_df = fights_df[["Date", "Event Link", "Fighter 1", "Fighter 1 Odds","Fighter 2", "Fighter 2 Odds"]]
470
+ fights_df = fights_df[
471
+ [
472
+ "Date",
473
+ "Event Link",
474
+ "Fighter 1",
475
+ "Fighter 1 Odds",
476
+ "Fighter 2",
477
+ "Fighter 2 Odds",
478
+ ]
479
+ ]
457
480
 
458
481
  cols_1 = [
459
482
  "Fighter 1",
@@ -470,8 +493,6 @@ mirrored_fights = ufc.mirror(
470
493
  cols_1,
471
494
  cols_2
472
495
  )
473
-
474
- mirrored_fights
475
496
  ```
476
497
 
477
498
  Using `in_order=True` produces rows in the following order:
@@ -483,7 +504,7 @@ Original Fight 2
483
504
  Mirrored Fight 2
484
505
  ...
485
506
  ```
486
- Using `in_order=True` produces rows in the following order:
507
+ Using `in_order=False` produces rows in the following order:
487
508
  ```text
488
509
  Original Fight 1
489
510
  Original Fight 2
@@ -494,6 +515,12 @@ Mirrored Fight 2
494
515
 
495
516
  This is useful when creating fighter-level datasets where each fight should be represented from both fighters' perspectives.
496
517
 
518
+ <br>
519
+
520
+ ---
521
+
522
+ <br>
523
+
497
524
  ## Train Test Split
498
525
  The function `train_test_split()` splits a UFC fight DataFrame into training and test sets while preserving the chronological order of the data. Unlike `sklearn.model_selection.train_test_split()`, this function does not randomly shuffle the data.
499
526
 
@@ -518,9 +545,10 @@ data = ufc.get_data()
518
545
 
519
546
  past_fights = data["past_fights"].copy()
520
547
  train, test = ufc.train_test_split(past_fights, 0.2)
521
- train
522
548
  ```
523
549
 
550
+ <br>
551
+
524
552
  ## Get Expanding Window
525
553
  The function `get_expanding_window()` creates multiple training and test sets for time-series cross-validation using an expanding training window.
526
554
 
@@ -567,6 +595,7 @@ train, test = ufc.train_test_split(past_fights, 0.2)
567
595
  folds = ufc.get_expanding_window(train, 10, 500)
568
596
  folds['Fold 2']['test']
569
597
  ```
598
+ <br>
570
599
 
571
600
  ## Plotting Accuracy
572
601
 
@@ -593,8 +622,127 @@ plot_test(accuracy, title="Test Accuracy")
593
622
  * `accuracy` — The accuracy.
594
623
  * `title` — The title of the plot.
595
624
 
625
+ <br>
626
+
596
627
  ## Expanding Window Baseline Example
597
628
 
629
+ In this example, we split the data into 80% for cross-validation and 20% for final testing. The cross-validation set is evaluated using a 15-fold expanding window approach. Using only betting odds, we convert the odds into implied probabilities and predict the fighter with the higher probability as the winner. The example also demonstrates how the model can be deployed.
630
+
631
+ <br>
632
+
633
+ ```python
634
+ import ufcdata as ufc
635
+ import numpy as np
636
+
637
+ data = ufc.get_data()
638
+
639
+ ## Select rows where outcome is W or L and both fighters have odds associated with the fights.
640
+ past_fights = data["past_fights"].copy()
641
+ past_fights = past_fights[~past_fights["Fighter 1 Outcome"].isin(["NC", "D"])]
642
+ past_fights = past_fights[past_fights[["Fighter 1 Odds", "Fighter 2 Odds"]].notna().all(axis=1)]
643
+
644
+ ## Convert odds to probabilities
645
+ past_fights["Fighter 1 Probability"] = (
646
+ past_fights["Fighter 1 Odds"]
647
+ .apply(ufc.convert_odds, probability=True)
648
+ )
649
+
650
+ past_fights["Fighter 2 Probability"] = (
651
+ past_fights["Fighter 2 Odds"]
652
+ .apply(ufc.convert_odds, probability=True)
653
+ )
654
+
655
+ ## Create the data for cross validation and final testing
656
+ folds = 15
657
+ fold_names = [f"Fold {i}" for i in range(1, folds + 1)]
658
+ cross_validation, final_test = ufc.train_test_split(past_fights, 0.2)
659
+ fifteen_fold_cv = ufc.get_expanding_window(cross_validation, folds, 500)
660
+
661
+ fold_accuracy = []
662
+ ## Loop through each fold to get the accuracy
663
+ for fold in fold_names:
664
+ test = fifteen_fold_cv[fold]['test']
665
+ test["Predicted Winner"] = np.where(
666
+ test["Fighter 1 Probability"] > test["Fighter 2 Probability"],
667
+ test["Fighter 1 Link"],
668
+ test["Fighter 2 Link"])
669
+ accuracy = float((test["Predicted Winner"] == test["Winner"]).mean())
670
+ fold_accuracy.append(accuracy)
671
+
672
+ ## Plot cross validation accuracy
673
+ ufc.plot_cv(fold_names, fold_accuracy)
674
+ ```
675
+ <br>
676
+
677
+ ![Cross Validation Plot](https://i.imgur.com/sjCpnRE.png)
678
+
679
+ <br>
680
+
681
+ ```python
682
+ ## Evaluate on final test set
683
+ final_test["Predicted Winner"] = np.where(
684
+ final_test["Fighter 1 Probability"] > final_test["Fighter 2 Probability"],
685
+ final_test["Fighter 1 Link"],
686
+ final_test["Fighter 2 Link"])
687
+ test_accuracy = float((final_test["Predicted Winner"] == final_test["Winner"]).mean())
688
+ ufc.plot_test(test_accuracy)
689
+ ```
690
+
691
+ <br>
692
+
693
+ ![Test Plot](https://i.imgur.com/5dHR59X.png)
694
+
695
+ <br>
696
+
697
+ ```python
698
+ ## Deploy the model
699
+ future_fights = data["future_fights"].copy()
700
+
701
+ future_fights["Fighter 1 Probability"] = (
702
+ future_fights["Fighter 1 Odds"]
703
+ .apply(ufc.convert_odds, probability=True)
704
+ )
705
+
706
+ future_fights["Fighter 2 Probability"] = (
707
+ future_fights["Fighter 2 Odds"]
708
+ .apply(ufc.convert_odds, probability=True)
709
+ )
710
+
711
+ future_fights["Predicted Winner"] = np.where(
712
+ future_fights["Fighter 1 Probability"] > future_fights["Fighter 2 Probability"],
713
+ future_fights["Fighter 1 Link"],
714
+ future_fights["Fighter 2 Link"])
715
+
716
+ future_fights["Predicted Winner"] = np.where(
717
+ future_fights["Predicted Winner"] == future_fights["Fighter 1 Link"],
718
+ future_fights["Fighter 1"],
719
+ future_fights["Fighter 2"]
720
+ )
721
+
722
+ future_fights = future_fights[
723
+ [
724
+ "Date",
725
+ "Event Link",
726
+ "Fight Link",
727
+ "Fighter 1",
728
+ "Fighter 2",
729
+ "Predicted Winner"
730
+ ]
731
+ ]
732
+
733
+ future_fights
734
+
735
+ ```
736
+ <br>
737
+
738
+ ![Deployed Predictions](https://i.imgur.com/jgZzNOV.png)
739
+
740
+ <br>
741
+
742
+ ---
743
+
744
+ <br>
745
+
598
746
  ## Online Sources
599
747
 
600
748
  Odds and birthplace data was obtained from https://www.tapology.com
@@ -603,12 +751,16 @@ Venue and attendance data was obtained from https://en.wikipedia.org/wiki/List_o
603
751
 
604
752
  All other data was obtained from http://ufcstats.com
605
753
 
754
+ <br>
755
+
606
756
  ## Discrepancies in Data
607
757
 
608
758
  UFCStats is treated as the authoritative source for UFCData. When discrepancies exist between UFCStats and other sources, such as Wikipedia or Tapology, the UFCStats data is used.
609
759
 
610
760
  The UFCStats completed events page serves as the authoritative source for event, fight, and round data. Individual fighter profiles may contain fights from organizations or events that are not included in the completed events database, including WEC, Strikeforce, and PRIDE. These events are therefore excluded from UFCData.
611
761
 
762
+ <br>
763
+
612
764
  ## Update Frequency
613
765
 
614
766
  The dataset is updated at the start of the scheduled broadcast time for each event. This update captures changes to betting odds, as well as any cancelled, postponed, or otherwise modified fights.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: UFCData
3
- Version: 0.7.0
3
+ Version: 0.7.2
4
4
  Summary: Obtain UFC data and functions to manipulate it
5
5
  Requires-Python: >=3.10
6
6
  License-File: LICENSE.md
@@ -0,0 +1,22 @@
1
+ [build-system]
2
+ requires = ["setuptools>=40.8.0", "wheel"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "UFCData"
7
+ version = "0.7.2"
8
+ description = "Obtain UFC data and functions to manipulate it"
9
+ requires-python = ">=3.10"
10
+
11
+ dependencies = [
12
+ "huggingface-hub",
13
+ "pandas",
14
+ "numpy",
15
+ "glicko2",
16
+ "rapidfuzz",
17
+ "matplotlib",
18
+ ]
19
+
20
+ [tool.setuptools.packages.find]
21
+ include = ["ufcdata*"]
22
+ exclude = ["scraper*", "UFCData_backup*"]
@@ -181,7 +181,7 @@ def get_fighter_statistic(fighter_link, fighter_bio, fights_df, rounds_df,
181
181
 
182
182
  record = get_fighter_history(fighter_link, fighter_bio, fights_df)
183
183
 
184
- date = ("2001-01-01")
184
+ date = ("2201-01-01")
185
185
 
186
186
  if len(record) == 0:
187
187
  columns = [
@@ -2,9 +2,9 @@ import re
2
2
  import pandas as pd
3
3
  import numpy as np
4
4
 
5
- def convert_odds(odds, probability=False):
5
+ def convert_odds(odds):
6
6
  """
7
- Convert American betting odds to a numeric value or implied probability.
7
+ Convert American betting odds to implied probability.
8
8
 
9
9
  Parameters
10
10
  ----------
@@ -12,23 +12,16 @@ def convert_odds(odds, probability=False):
12
12
  American betting odds. The function extracts the first signed or
13
13
  unsigned integer from the input. Missing values return None.
14
14
 
15
- probability : bool, default=False
16
- If False, return the extracted American odds as an integer.
17
- If True, return the implied probability as a decimal.
18
-
19
15
  Returns
20
16
  -------
21
- int or float or None
22
- The extracted American odds if probability=False, otherwise the
23
- implied probability. Returns None for missing values.
17
+ float or None
18
+ The implied probability as a decimal. Returns None for missing values.
24
19
 
25
20
  Examples
26
21
  --------
27
22
  >>> convert_odds("+120")
28
- 120
23
+ 0.45454545454545453
29
24
  >>> convert_odds("-400")
30
- -400
31
- >>> convert_odds("-400", probability=True)
32
25
  0.8
33
26
  """
34
27
  if pd.isna(odds):
@@ -36,9 +29,6 @@ def convert_odds(odds, probability=False):
36
29
 
37
30
  odds = int(re.search(r'[+-]?\d+', str(odds)).group())
38
31
 
39
- if not probability:
40
- return odds
41
-
42
32
  if odds < 0:
43
33
  return -odds / (-odds + 100)
44
34
  else:
@@ -1,14 +0,0 @@
1
- [project]
2
- name = "UFCData"
3
- version = "0.7.0"
4
- description = "Obtain UFC data and functions to manipulate it"
5
- requires-python = ">=3.10"
6
-
7
- dependencies = [
8
- "huggingface-hub",
9
- "pandas",
10
- "numpy",
11
- "glicko2",
12
- "rapidfuzz",
13
- "matplotlib",
14
- ]
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes
File without changes