Assignment1 Assignment2 Assignment3 Assignment4 Assignment5 Assignment6 Assignment7 Assignment8

Assignment 6

Predicting the outcome of a chess game

1.The data was taken from Kaggle
2.Attributes used : respective ratings of the player, number of turns, kind of openings
3.Various classification model were tried to predict the outcomes
4. The best model came out to be the Fine Tree model
5. The prediction made was which color will have a higher chances of winning

Motivation

Lots of information is contained within a single chess game, let alone a full dataset of multiple games. It is primarily a game of patterns, and data science is all about detecting patterns in data, which is why chess has been one of the most invested in areas of AI in the past. This dataset collects all of the information available from 20,000 games and presents it in a format that is easy to process for analysis of, for example, what allows a player to win as black or white, how much meta (out-of-game) factors affect a game, the relationship between openings and victory for black and white and more. And hence it looked an interesting problem to look into.

Source of data

Link

Pre processing script used

From the data the additional information like the player id, Opening echo, moves, increment code, time, victory status were ignored because they didnt have any correlation. Apart from this the data was already pre processed.

The Machine Learning models used

Fine Tree Model - 60.2%
Medium Tree Model-59.8%
Coarse Tree Model - 56.5%
Various kind of KNN(K nearest neighbours) models - 60.9%
The best accuracy was with Fine Tree model and is familiar to me as well and hence it was considered to predict the outcomes. The low accuracy can be attributed to overfitting, because of too much data points ( more than 20000 points)

Results

scatter plot with same attributes so that dependencey with particular attribute can be identified confusion matrix scatter plot

Blue - Person moving with black is the winner; Red - draw; Orange - white is the winner

1. For larger number of turns, more number of draws were there and for lesser number of turns we mostly get a winner
2. The winner was more likely the player with higher rating as is evident from the scatter plot
3. From confusion matrix it is evident that the model's prediction have been good not very accurate but most of the predictions and the actual results matches

Improvement

If the overfitting could be avoided then the results might improve