League of Legends
Analysis for Game Result Prediction
This is DSC 80 Project 4 where we clean, explore, and make predictions upon, League of Legends game data. This website stands as a report of our findings.
Names: Rihui Ling and Minghan Wu
Introduction
Our dataset is on all the professional League of Legends games between 2014 and 2024 (inclusive).
The dataset contains one row per game per team (so two rows per game). The
game data is comprehensive of most aspects of the game. Our cleaned game data has 148989 rows × 26 columns.
In this report, we discuss one question, and one prediction problem.
Question to Answer:
In the history of LoL, the side of team (blue or red) is often believed to have an influence on the team winning rate, with the blue side having a greater chance of winning. If this is true, then the result of game largely depends on luck – the team on the blue side has a significant advantage – and hence the outcome of one game does not necessarily reflect the strength of a team. For this reason, we put forth this question:
Will the side of team by a team in a game affect the team kills?
-
\(H_0\): The blue side has the same expected team kills as the red side.
-
\(H_1\): The blue side has higher expected team kills than the red side.
After that, we will come up with model(s) for the following prediction problem:
Prediction Problem
- Predict whether a team is winning/losing a game, given other data.
While much of the data we have is only available post game, such a model can indicate what factors are important for winning a LoL game.
Data
Our raw data has 897996 rows × 130 columns. Most of the columns are of little or no help for our analysis.
Here lists the columns we decide to keep along with their description:
- Basic game stats:
gameid: the game idgame: the ordinal number of the game in the competitionpatchthe major patch numbergamelengththe time span of the game in minutes
- Team game stats:
side: the side of the team, blue or redresult: result of the team in the gamekills: the total kills of the team in the gamedeaths: the total deaths of the team in the gameassists: the total assists of the team in the gamefirstblood: whether the team got first blooddamagetochampions: the total damage made to the enemy teamearnedgold: the total number of gold earned by the teamfirsttower: whether the team got the first towertowers: the number of towers got by the teamopp_towers: the number of towers got by the enemy teamfirstmidtower: whether the team got the first middle towerfirsttothreetowers: whether the team got three towers before the enemy team
- General game resources:
firstdragon: whether the team got the first dragondragons: the total number of dragons got by the teamopp_dragons: the total number of dragons got by the enemy teamfirstherald: whether the team got the first heraldelders: the total number of elders got by the teamopp_elders: the total number of elders got by the enemy teamfirstbaron: whether the team got the first baronbarons: the total number of barons got by the teamopp_barons: the total number of barons got by the enemy team
Data Cleaning and Exploratory Data Analysis
Data Cleaning
1 We only consider team data, so drop all player rows, only keeping the rows where position is team
2 Drop rows that the earned gold is less than or equal to 0, since it should only contain non-negative value and value 0 indicates very short game length.
3 Convert patch to major patch
4 Drop all rows where gamelength is greater than 2 hrs (since the longest game in the history of LOL is about 1h30min), convert the unit of gamelength from s to min
5 Keep only the columns useful for our analysis
6 Convert all binary encoded columns to bool
7 Convert all numerical column with integral values to int
We follow these steps so only team data is kept, which all alysis is based on team. Moreover, remove those extreme value will improve the significance of the analysis.
This table shows the first several rows of the dataframe after data cleaning:
** Cleaned DataFrame Shape: 148989 rows × 26 columns
Univariate Analysis
Here is a histogram showing the distribution of gamelength:
We can see that the data is roughly normally distributed, centering around 32.5 min. Note that there are outliers on both ends of the graph (even after we remove the extreme values), although their density is too small to be visible.
Bivariate Analysis
Here is a graph showing the distributions of gold earned by the team in the game conditional on game results (win or lose):
Both distributions are roughly normally distributed, with the win distribution more concentrated than the lose distribution, and with higher mean.
Interesting Aggregates
The following pivot table shows the average team kills for lost and won teams for different years.
We can see that for each year, the difference between average team kills for won teams and lost teams are around 10 (so kills and result are strongly related). Also, the average team kills for both won teams and lost teams increased significantly since 2021.
Missingness Mechanisms
In this section, we analyze the missing mechanisms of several of the columns. To give an overview, here is a DataFrame showing how many rows are missing for each column:
Not Missing at Random (NMAR)
We believe the missingness of game is NMAR. Upon researching online, we discover that the missingness results largely from the fact that there is only one game in the split in that league and year.
As such, the missing mechanism of game is NMAR, with value 1 (representing the first game) being significantly more prone to be missing than other values.
Missing at Random (MAR)
We believe the missingness of elders and opp_elders is not NMAR because there is no reason why the missingness of elders is dependent on the (missing) values themselves.
However, we believe the missingness of elders and opp_elders is dependent on patch. To confirm this idea, we run a permutation test on the two columns (using Total Variation Distance (tvd) as statistic).
This is the observed conditional distribution:
This is the DataFrame showing the same distribution:
After running the permutation and computing Total Variation Distance (TVD) 500 times, this is the distribution we get:
We see that the observed tvd \(0.615131\) is significantly greater than all tvd values from the distribution. The p-value for the permutation test is \(0.0\).
Hence, we conclude that the missingness of elders is dependent on patch, making the former missing at random.
After running the same test on the missingness of opp_elders, we find the same pattern (which is expected because it is inherently in pairs with elders)
Missing Completely at Random (MCAR)
We believe the missing mechanism of barons and opp_barons is completely at random. To confirm this, we run permutation tests against each every one of other columns.
For numerical columns, we compute the absolute differences between the means of the missing group and the non-missing group;
For categorical columns, we compute the tvds between the missing group and the non-missing group.
Here is the DataFrame showing our result for barons:
From this, we can see that the missingness of barons is independent of the values of other columns, with 95% confidence level.
(Some p-values are close to 1, this is because the number of missing barons is very few.)
After running the same test on the missingness of opp_barons, we got the same result.
Handling Missingness (Imputation)
-
Perform probabilistic imputation on missing
eldersandopp_eldersconditional onpatch. We usenp.random.choiceto randomly select \(n\)elders/opp_eldersvalues from the rows with the samepatch, where \(n\) is the number ofelders/opp_eldersmissing corresponding to thatpatch(we regard na value inpatchas a valid category). Note that for those patches don’t have elders, the non-missing values are all 0, and hence the imputed values must be 0. -
Listwise delete rows where
baronsandopp_baronsis missing (delete all rows wherebaronsandopp_baronsis missing), sincebaronsandopp_baronsis missing completely at random.
Hypothesis Testing
As mentioned in the introduction, in the history of LoL, the side of the team (blue or red) is often believed to have an influence on the team winning rate, with the blue side having a greater chance of winning. In this section, we want to test out will the kills, which is high correlated with win rate, for each side is having significant difference.
Question: Will the of side of the team in a game affect the team kills?
-
\(H_0\): The blue side has the same expected team kills as the red side.
-
\(H_1\): The blue side has higher expected team kills than the red side.
-
Test statistic: average kills blue - average kills red (Since we are perform a one side hypothesis test)
-
Significance level: 5%
Here is a graph showing the overlaying distributions of red-side kills and blue-side kills:
This dataframe shows the observed data:
We test the two hypotheses by conducting a permutation test by permuting kills 1000 times and compute the test statistic .
Below is a graph showing the distribution of the test statistic according to \(H_0\) and our observed statistic.
Since p-value we obtain is 0.0 and \(0.0<0.05\), we reject \(H_0\). This test makes us conclude that more likely than not, the team being on the blue side does increase the expected kills of the team.
Framing a Prediction Problem
Prediction Problem: We try to obtain a model that predicts whether a team is winning/losing a game, given the data we kept.
This is a binary classification problem.
We will use (test set) accuracy to evaluate our model, since
- It is a more direct metric for the success of the model
- We do not have a preference towards false positive / false negative
- The
resultwe get after cleaning and imputation has about equal numbers ofTrueandFalse(the percentage of winning is \(50.01%\))
Baseline Model
Our baseline model uses logistic regression to predict whether the team will win or lose given the 22 features we select.
Here is a dataframe showing the features we use in our baseline model, separated by category (nominal, ordinal, numerical).
For this model, all features come from the cleaned dataframe after imputation.
Here, we describe what transformation we have applied to each category:
Nominal
We have 8 nominal columns.
We one-hot-encode every categorical column (each time, drop one of the columns generated to avoid colinearity).
Numerical
We have 14 numerical columns.
We standardize every numerical column.
That is, for every column \(X=(X_1, X_2, ..., X_n)\), for every \(i=1, 2, ..., n\), we let \(X_{std, i}=\frac{X_i-\bar{X}}{SD_X}\).
Ordinal
We do not have any ordinal feature in our selection, since no categorical column has more than 2 values that have an inherent order i.e. exchanging the labels makes no difference.
Model Parameters
For this model, we use default parameters of sklearn.LogisticRegression.
- Penalty:
'l2' - tol (tolerance for stopping iteration):
1e-4 - max_iter:
100 - solver:
lbfgs
Model Result
The following dataframe shows the model accuracy:
The following graph is the confusion matrix on the test set:
We believe this model is effective, based on the high accuracy it gets.
Final Model
In this model, we have added derivative features to the dataframe, together with the original 22 features. The final model has 30 features in total.
Here are the 8 new features:
kda: Calculated by \(\frac{kills+0.5\times assists}{deaths}\) (If for a rowdeathis 0, we setdeathto 1). This is a commonly used measure for team performance. It takes into account the ratio of kills and assists to deaths.soulopp_soul: Binary indicator of whetherdragons/opp_dragonsis greater than or equal to 4. This is important because in later patch, a team gets dragon soul buff once they get 4 dragons.kills_per_mindeaths_per_minassists_per_mindamagetochampions_per_minearnedgold_per_min: Those columns divided bygamelength. This can be helpful because they indicate the rate at which the data changes in the game.
For the final model, we decide to stick to logistic regression.
In order for the model to be fairly comparable to the baseline model, we use the same training set and test set as before.
Best Hypermeters
We want to optimize the following hypermeters:
max_iterMaximum number of iterations taken for the solvers to converge.
- Range: \(\{100, 150, 250, 300\}\)
solverAlgorithm to use in the optimization problem.
- Range: {“lbfgs”, “libliear”, “newton-cg”, “sag”, “saga”}
Since tol (Tolerance for stopping criteria) is highly related to
max_iter, we do not optimize it. Instead, we keep it as default
\(10^{-4}\)
(or 1e^-4).
To determine the best combination, we use GridSearchCV to perform a grid search (comparing all combinations of max_iter and solver) with 5-fold cross-validations on the training set.
This is the result of GridSearchCV:
From this process, we conclude that the best hypermeters combination is
-
max_iter = 100 -
solver = "newton-cg".
Model Result
The following dataframe shows the model accuracy:
Here shows the confusion matrix on the test set of the optimal model that we found:
Improvement of Final Model
Here is a summary of the model accuracy of the baseline model and the final model. There is a slight improvement compared to the baseline model (0.0442% improved on predict test dataset).
The two confusion matrices for two models:
Fairness Analysis
Could our model perform worse for blue or red side due to the fact that they start in slightly different positions?
In this section, we want to explore the consistency of the performance of our model for red and blue sides. We want to find out whether the final model achieve accuracy parity for blue and red sides. i.e. is the difference in accuracy due to randomness?
Two Groups: blue side and red side
\(H_0\): Our model is fair. That is, the accuracy of the final model for blue side is the same as the accuracy for red side.
\(H_1\): The accuracy of the final model for blue side is smaller than the accuracy for the red side.
Test Statistic: \(accuracy(blue)-accuracy(red)\)
Significance Level: 5%
Evaluation Metric: accuracy
We run a permutation test to determine which hypothesis is correct:
We compute the observed difference, and simulate the \(H_0\) distribution by permuting side 1000 times and compute the difference between the accuracy for “blue” side and the accuracy for “red” side.
This dataframe shows the observed statistic:
This graph shows the result of the permutation test:
Since the p-value is 0.294 and \(0.294>0.05\), we fail to reject \(H_0\) at 5% significance level. The difference in observed accuracy is not statistically significant to suggest that our model works better on blue side than on red side.