01
“If I were not a physicist, I would probably be a musician. I often think in music.
I live my daydreams in music. I see my life in terms of music.” – Albert Einstein
I might not be a physicist like Mr Einstein, but I agree wholeheartedly with his views on music!
I can not recall a single day when my music player wasn't set up. We’ve always dreamed of composing
music but couldn’t quite get the hang of instruments. That was until I came across machine learning.
Using certain techniques and frameworks, I was able to compose my own original music score without
really knowing any music theory! And this forms the basics of my project Automatic Music Generation.
I also selected this topic to work upon, because none of my equipments related to arduino were working
other than a buzzer, since I left most of them in my hostel room, and couldn't order it because of non-availability
of delivery services in my area.

02
The Automatic Music Generation is a process of composing a short piece of music with minimum
human intervention.
Music is essentially composed of Notes and Chords. From the perspective of a piano instrument, we have:
Note: The sound produced by a single key is called a note
Chords: The sound produced by 2 or more keys simultaneously is called a chord. Generally, most chords contain at least 3 key sounds
Octave: A repeated pattern is called an octave. Each octave contains 7 white and 5 black keys
03
Deep Learning is a field of Machine Learning which is inspired by a neural structure. These networks extract the features automatically from the dataset and are capable of learning any non-linear function. That’s why Neural Networks are called as Universal Functional Approximators.
WaveNet is a Deep Learning-based generative model for raw audio developed by Google DeepMind.
The main objective of WaveNet is to generate new samples from the original distribution of the data. Hence, it is known as a Generative Model.
THE TRAINING PHASE
Input to the WaveNet
WaveNet takes the chunk of a raw audio wave as an input. Raw audio wave refers to the representation of a wave in the time series domain.
In the time-series domain, an audio wave is represented in the form of amplitude values which are recorded at different intervals of time:
Output from the WaveNet
Given the sequence of the amplitude values, WaveNet tries to predict the successive amplitude value.
Let’s understand this with the help of an example. Consider an audio wave of 5 seconds with a sampling rate of 16,000 (that is 16,000 samples per second). Now, we have 80,000 samples recorded at different intervals for 5 seconds. Let’s break the audio into chunks of equal size, say 1024 (which is a hyperparameter).
We can follow a similar procedure for the rest of the chunks.
We can infer from the above that the output of every chunk depends only on the past information ( i.e. previous timesteps) but not on the future timesteps. Hence, this task is known as Autoregressive task and the model is known as an Autoregressive model.
Inference phase
In the inference phase, we will try to generate new samples. The steps of which are:
1. Select a random array of sample values as a starting point to model
2. Now, the model outputs the probability distribution over all the samples
3. Choose the value with the maximum probability and append it to an array of samples
4. Delete the first element and pass as an input for the next iteration
5. Repeat steps 2 and 4 for a certain number of iterations
Understanding the WaveNet Architecture
The building blocks of WaveNet are Causal Dilated 1D Convolution layers. One of the main reasons for using convolution is to extract the features from an input.
Convolution is a mathematical operation that combines 2 functions. In the case of image processing, convolution is a linear combination of certain parts of an image with the kernel.
What is 1D Convolution?
The objective of 1D convolution is similar to the Long Short Term Memory model. It is used to solve similar tasks to those of LSTM. In 1D convolution, a kernel or a filter moves along only one direction.
The output of convolution depends upon the size of the kernel, input shape, type of padding, and stride. Now, I will walk you through different types of padding for understanding the importance of using Dilated Causal 1D Convolution layers.
When we set the padding valid, the input and output sequences vary in length. The length of an output is less than an input:
When we set the padding to same, zeroes are padded on either side of the input sequence to make the length of input and output equal:
Pros of 1D Convolution specific to this project
• Captures the sequential information present in the input sequence.
• Training is much faster compared to GRU or LSTM because of the absence of recurrent connections.
Cons of 1D Convolution specific to this project
• When padding is set to the same, output at timestep t is convolved with the previous t-1 and future timesteps t+1 too. Hence, it violates the Autoregressive principle
• When padding is set to valid, input and output sequences vary in length which is required for computing residual connections
What is 1D Causal Convolution?
This is defined as convolutions where output at time t is convolved only with elements from time t and earlier in the previous layer.
In simpler terms, normal and causal convolutions differ only in padding. In causal convolution, zeroes are added to the left of the input sequence to preserve the principle of autoregressive.
Pros of Causal 1D convolution:
• Causal convolution does not take into account the future timesteps which is a criterion for building a Generative model
Cons of Causal 1D convolution:
• Causal convolution cannot look back into the past or the timesteps that occurred earlier in the sequence. Hence, causal convolution has a very low receptive field. The receptive field of a network refers to the number of inputs influencing an output:
As you can see here, the output is influenced by only 5 inputs. Hence, the Receptive field of the network is 5, which is very low. The receptive field of a network can also be increased by adding kernels of large sizes but keep in mind that the computational complexity increases.
This drives us to the awesome concept of the Dilated 1D Causal Convolution.
What is Dilated 1D Causal Convolution?
A Causal 1D convolution layer with the holes or spaces in between the values of a kernel is known as Dilated 1D convolution.
The number of spaces to be added is given by the dilation rate. It defines the reception field of a network. A kernel of size k and dilation rate d has d-1 holes in between every value in kernel k.
As you can see here, convolving a 3 * 3 kernel over a 7 * 7 input with dilation rate 2 has a reception field of 5 * 5.
Pros of Dilated 1D Causal Convolution:
• The dilated 1D convolution network increases the receptive field by exponentially increasing the dilation rate at every hidden layer:
As you can see here, the output is influenced by all the inputs. Hence, the receptive field of the network is 16.
The Residual block of Wavenet
A building block contains Residual and Skip connections which are just added to speed up the convergence of the model
The Workflow of WaveNet:
• Input is fed into a causal 1D convolution
• The output is then fed to 2 different dilated 1D convolution layers with sigmoid and tanh activations
• The element-wise multiplication of 2 different activation values results in a skip connection
• And the element-wise addition of a skip connection and output of causal 1D results in the residual
04
I downloaded and combined multiple classical music files of a digital piano from numerous resources. The final dataset is here.
Importing Libraries
Music 21 is a Python library developed by MIT for understanding music data. MIDI is a standard format for storing music files. MIDI stands for Musical Instrument Digital Interface. MIDI files contain the instructions rather than the actual audio. Hence, it occupies very little memory. That’s why it is usually preferred while transferring files.
Reading Musical Files
Let’s define a function straight away for reading the MIDI files. It returns the array of notes and chords present in the musical file.
Now, we will load the MIDI files into our environment
We will now explore the dataset and understand it in detail.
Output: 304
As you can see here, no. of unique notes is 304. Now, let us see the distribution of the notes.
Output
From the above plot, we can infer that most of the notes have a very low frequency. So, let us keep the top frequent notes and ignore the low-frequency ones. Here, I am defining the threshold as 50. Nevertheless, the parameter can be changed.
Output: 167
As you can see here, no. of frequently occurring notes is around 170. Now, let us prepare new musical files which contain only the top frequent notes
Preparing Data
Preparing the input and output sequences
Now, we will assign a unique integer to every note and prepare the integer sequences for input data.
Similarly, we will prepare the integer sequences for output data as well and preserve 80% of the data for training and the rest 20% for the testing:
Model Building
Simplifying the architecture of the WaveNet without adding residual and skip connections since the role of these layers is to improve the faster convergence (and WaveNet takes raw audio wave as input). But in our case, the input would be a set of nodes and chords since we are generating music:
Define the callback to save the best model during training and train the model with a batch size of 128 for 50 epochs. And saving the best model for using further.
Its time to compose our own music now. We will follow the steps mentioned under the inference phase for the predictions. Then we will convert the integers back into the notes.
And at the end, we will convert back the predictions into a MIDI file, and save it into a musical file.
04
• As the size of the training dataset is small, we can fine-tune a pre-trained model to build a robust system.
• Collecting as much as training data as possible, since the deep learning model generalizes well on the larger datasets.
05
Now, since we have generated the MIDI file, we can also play it using a buzzer and an
Arduino. But first we will have to convert the MIDI file into notes which arduino can
play. We will accomplish this using this website: MIDI to Arduino Notes
And then, we will play the notes on the arduino and the buzzer.
I built an app using MIT App Inventor which connects using HC-05 bluetooth module and turns the
music on and off, the code of which is uploaded in the code section below.

The Video.
This is the end of this project.
06
The repository for the model and the code for arduino are as follows:
1. Model Github Repository
2. Arduino Code
3. App Inventor File