AIL 722/7022: Reinforcement Learning (Fall 2025)

Course Description and Objectives

In this course, we will develop the knowledge and skills to understand, implement and apply recent deep reinforcement learning techniques. This course is intended for graduate and advanced undergraduate students.

Specifically, we set the following course objectives:

  1. Formulate sequential decision making problems as Markov Decision Processes and map them to suitable RL algorithms.
  2. Use deep RL algorithms on standard benchmark environments.
  3. Deploy RL algorithms in various domains in conjunction with other ML techniques.

Prerequisites:

  1. Theory: Probability and statistics; Algorithms; Exposure to AI and/or ML will be helpful
  2. Implementation: Fluency in a high level programming language; Familiarity with installing and working with Github packages; Experience with using high-performance compute services will be helpful

Logistics

Time and Location:

  • Slot: AC
  • Class Timings: Tue/Fri 2-3:20 pm
  • Venue: LH313.5

  • Office hours
    • Timing: By appointment
    • Venue: 513-D, Academic Complex East

Communication:

We will use piazza as the forum for students to ask questions about both the course material as well as logistics.

Joining Link

Teaching Team:

Course Contents

Below is a tentative list of topics:

  1. Markov Decision Processes
  2. Dynamic Programming; Approximations
  3. Monte Carlo Prediction and Control
  4. Temporal-Difference Methods
  5. Exploration-Exploitation Tradeoff
  6. Policy Gradients
  7. Imitation learning
  8. Offline RL

Textbooks

  1. Introduction to Reinforcement Learning, Sutton and Barto
  2. Algorithms for Optimisation, Kochenderfer and Wheeler
  3. Algorithms for Decision Making, Kochenderfer, Wheeler and Wray
  4. Planning with Markov Decision Processes, Mausam and Kolobov

Announcements

Announcements will appear here.

Schedule

The weekly schedule will be updated with lecture slides and pointers to additional references as we progress through the course.

Monday Tuesday Wednesday Thursday Friday Saturday Sunday
Week 1 July 21 July 22 July 23 July 24 July 25 July 26 July 27
Course overview Lecture 1:

Course overview
Slides
Week 2 July 28 July 29 July 30 July 31 Aug 1 Aug 2 Aug 3
Stochastic Sequential Models Lecture 2:

Hidden Markov Models
Slides

References:
Jurafsky
Lecture 3:

Markov Decision Processes
Slides

References:
S&B (chp 3)

Additional Readings:
Bellman,1957
Week 3 Aug 4 Aug 5 Aug 6 Aug 7 Aug 8 Aug 9 Aug 10
MDPs and Value Functions Lecture 4:

Value Functions
Slides

References:
S&B (chp 3)
Lecture 5:

MDP Formalisation
Slides

References:
M&K (chp 2)
Week 4 Aug 11 Aug 12 Aug 13 Aug 14 Aug 15 Aug 16 Aug 17
Policy Evaluation Lecture 6:

Dynamic Programming
Slides

References:
S&B (chp 4)

Additional Readings:
Minsky,1961
Lecture 7:

Monte Carlo
Slides

References:
S&B (chp 5)

Additional Readings:
Barto & Duff,1994
Michie & Chambers,1968
Widrow, Gupta & Maitra,1973
Week 5 Aug 18 Aug 19 Aug 20 Aug 21 Aug 22 Aug 23 Aug 24
Policy Improvement Lecture 8:

Policy Improvement
Slides

References:
S&B (chp 4)

Additional Readings:
Puterman & Shin,1978
Lecture 9:

Policy Iteration and Value Iteration
Slides

References:
S&B (chp 4)

Additional Readings:
Connell,1989
Week 6 Aug 25 Aug 26 Aug 27 Aug 28 Aug 29 Aug 30 Aug 31
Model-Free Prediction and Control Lecture 10:

Temporal Difference Prediction
Slides

References:
S&B (chp 5&6)

Additional Readings:
Dayan,1992
Jaakkola, Jordan & Singh,1993
Sutton,1988
Tsitsiklis,1994
Lecture 11:

Batch TD and MC; Monte Carlo Control
Slides

References:
S&B (chp 5&6)

Additional Readings:
Singh & Sutton,1996
Week 7 Sep 1 Sep 2 Sep 3 Sep 4 Sep 5 Sep 6 Sep 7
Monte Carlo Control Lecture 12:

Monte Carlo Control
Slides

References:
S&B (chp 5)

Additional Readings:
Singh & Sutton,1996
Michie & Chambers,1968
Week 8 Sep 8 Sep 9 Sep 10 Sep 11 Sep 12 Sep 13 Sep 14
TD Control Lecture 13:

TD Control
Slides

References:
S&B (chp 6)
Week 9 Sep 15 Sep 16 Sep 17 Sep 18 Sep 19 Sep 20 Sep 21
Week 10 Sep 22 Sep 23 Sep 24 Sep 25 Sep 26 Sep 27 Sep 28
Exploration Lecture 14:

Exploration
Slides

References:
S&B (chp 2)
Week 11 Sep 29 Sep 30 Oct 1 Oct 2 Oct 3 Oct 4 Oct 5
Week 12 Oct 6 Oct 7 Oct 8 Oct 9 Oct 10 Oct 11 Oct 12
Approximating Value Functions Lecture:

HPC Intro, PyTorch Tutorial
Slides
Lecture 15:

Approximating value functions
Slides

References:
S&B (chp 9)
Week 13 Oct 13 Oct 14 Oct 15 Oct 16 Oct 17 Oct 18 Oct 19
DQN Lecture 16:

Deep Q Networks
Slides

References:
Mnih, 2013
Lecture 17:

Double Q learning; PG Intro
Slides

References:
Optimizer's Curse
van Hasselt, 2010
Week 14 Oct 20 Oct 21 Oct 22 Oct 23 Oct 24 Oct 25 Oct 26
Policy Optimisation Lecture 18:

Policy Optimisation
Slides
Week 15 Oct 27 Oct 28 Oct 29 Oct 30 Oct 31 Nov 1 Nov 2
Policy Gradients Lecture 19:

Policy Gradients
Slides

References:
S&B (chp 13)
Lecture 20:

Actor-Critic Methods
Slides

References:
A3C, 2015

Additional Readings:
Natural Actor-Critic, 2009
Week 16 Nov 3 Nov 4 Nov 5 Nov 6 Nov 7 Nov 8 Nov 9
Policy Gradient Extensions Lecture 21:

Extensions to Policy Search
Slides
Lecture 22:

Surrogate Advantage Objective


References:
TRPO, 2015
PPO, 2017

Additional Readings:
NPG
Week 17 Nov 10 Nov 11 Nov 12 Nov 13 Nov 14 Nov 15 Nov 16
Off-Policy and Max Entropy Lecture 23:

Model-Based RL; Max-Entropy
Slides

References:
S&B (chp 8)
Lecture 24:

Soft VI; Soft PI; SAC
Slides

References:
Soft Q Learning, 2017
Soft Actor Critic, 2018

Grading Policy

Grading policy (tentative, upto ±15% on each component):

  1. Programming Assignments: 60%
  2. Exams: 40%
  3. Audit policy: Only based on exams: To get an audit pass, you should obtain at least 50% absolute score on each and every exam.