1.The data was taken from Kaggle
2.Attributes used : respective ratings of the player, number of turns, kind of openings
3.Various classification model were tried to predict the outcomes
4. The best model came out to be the Fine Tree model
5. The prediction made was which color will have a higher chances of winning
Lots of information is contained within a single chess game, let alone a full dataset of multiple games. It is primarily a game of patterns, and data science is all about detecting patterns in data, which is why chess has been one of the most invested in areas of AI in the past. This dataset collects all of the information available from 20,000 games and presents it in a format that is easy to process for analysis of, for example, what allows a player to win as black or white, how much meta (out-of-game) factors affect a game, the relationship between openings and victory for black and white and more. And hence it looked an interesting problem to look into.
From the data the additional information like the player id, Opening echo, moves, increment code, time, victory status were ignored because they didnt have any correlation. Apart from this the data was already pre processed.
Fine Tree Model - 60.2%
Medium Tree Model-59.8%
Coarse Tree Model - 56.5%
Various kind of KNN(K nearest neighbours) models - 60.9%
The best accuracy was with Fine Tree model and is familiar to me as well and hence it was considered to predict the outcomes. The low accuracy can be attributed to overfitting, because of too much data points ( more than 20000 points)
Blue - Person moving with black is the winner; Red - draw; Orange - white is the winner
1. For larger number of turns, more number of draws were there and for lesser number of turns we mostly get a winner
2. The winner was more likely the player with higher rating as is evident from the scatter plot
3. From confusion matrix it is evident that the model's prediction have been good not very accurate but most of the predictions and the actual results matches
If the overfitting could be avoided then the results might improve