AIL 8027/841: Advanced Reinforcement Learning (Spring 2026)

Course Description and Objectives

In this course, we will develop the knowledge and skills to understand, implement and apply recent deep reinforcement learning techniques. This course is intended for graduate and advanced undergraduate students.

Specifically, we set the following course objectives:

  1. Learn about techniques to design and analyse RL algorithms
  2. Learn about the foundational principles behind recent advances in deep RL
  3. Gain awareness of the frontier research areas in deep RL

Prerequisites:

  1. Theory: Probability and statistics; Algorithms; Exposure to AI and/or ML will be helpful
  2. Implemenation: Fluency in a high level programming language; Familiarity with installing and working with Github packages; Experience with using high-performance compute services will be helpful

Logistics

Time and Location:

  • Slot: C
  • Class Timings: Tue, Wed, Fri 8 am
  • Venue: LHC 518

  • Office hours
    • Timing: By appointment
    • Venue: 513-D, Academic Complex East

Communication:

We will use piazza as the forum for students to ask questions about both the course material as well as logistics. Joining link to be shared.

Joining Link

Teaching Team:

Course Contents

Below is a tentative list of topics:

  1. Analysis of tabular RL algorithms
  2. Approximations
  3. Policy Search
  4. Probabilistic Graphical models and RL
  5. Safe RL
  6. Offline RL

Textbooks

  1. Introduction to Reinforcement Learning, Sutton and Barto
  2. Algorithms for Optimisation, Kochenderfer and Wheeler
  3. Algorithms for Decision Making, Kochenderfer, Wheeler and Wray
  4. Numerical Optimisation, Nocedal and Wright
  5. Neuro-Dynamic Programming, Bertsekas and Tsitsiklis
  6. Introduction to Markov Decision Processes (under preparation), Puterman

Announcements

Announcements will appear here.

Schedule

The weekly schedule will be updated with lecture slides and pointers to additional references as we progress through the course.

Monday Tuesday Wednesday Thursday Friday Saturday Sunday
Week 1 Jan 5 Jan 6 Jan 7 Jan 8 Jan 9 Jan 10 Jan 11
Course overview, Bellman Equations

Reference:
Puterman Chapter 5
Lecture 1:

Course Overview
Slides
Lecture 2:

Bellman Equations
Slides

References:
S&B
Lecture 3:

Iterative Policy Evaluation
Slides
Lecture 4:

Contraction Proof, Bellman Optimality Intro
Week 2 Jan 12 Jan 13 Jan 14 Jan 15 Jan 16 Jan 17 Jan 18
Contraction Mapping

Reference:
Puterman Chapter 5
Lecture 5:

Bellman Optimality Eqn Proof
Week 3 Jan 19 Jan 20 Jan 21 Jan 22 Jan 23 Jan 24 Jan 25
Value Iteration Analysis

Reference:
Puterman Chapter 5
Lecture 6:

Value Iteration
Lecture 7:

Bellman Optimality Contraction
Lecture 8:

Quality of Policy in VI, PI Intro
Week 4 Jan 26 Jan 27 Jan 28 Jan 29 Jan 30 Jan 31 Feb 1
Policy Iteration, Approximate PI

References:
Bertsekas, 2011
Scherrer
Lecture 9:

Policy Iteration Analysis
Lecture 10:

Approximate Policy Iteration Intro
Lecture 11:

API: Dataset Generation
Week 5 Feb 2 Feb 3 Feb 4 Feb 5 Feb 6 Feb 7 Feb 8
Searching in Policy Space Lecture 12:

API: No Monotonicity

References:
Pirotta, 2013
Wagner, 2011
Lecture 13:

Expected Advantage Objective

References:
Kakade, 2002
Lecture 14:

Performance Difference Lemma
Week 6 Feb 9 Feb 10 Feb 11 Feb 12 Feb 13 Feb 14 Feb 15
Performance Difference Lemma Lecture 15:

PDL proof

Quiz Lecture 16:

CPI Analysis

References:
Kakade, ICML 2002
Week 7 Feb 16 Feb 17 Feb 18 Feb 19 Feb 20 Feb 21 Feb 22
Conservative Policy Iteration

References:
Kakade, ICML 2002
Lecture 17:

CPI monotonicity
Lecture18:

CPI from PDL
Week 8 Feb 23 Feb 24 Feb 25 Feb 26 Feb 27 Feb 28 Mar 1
Week 9 Mar 2 Mar 3 Mar 4 Mar 5 Mar 6 Mar 7 Mar 8
Week 10 Mar 9 Mar 10 Mar 11 Mar 12 Mar 13 Mar 14 Mar 15
Policy Optimisation:

Trust-region methods
Lecture 19:

Parameterised policies; KL divergence constraint

References:
Kakade, NeurIPS 2001
Lecture 20:

Optimisation; Linear Objective Quadratic Constraint

Lecture 21:

TRPO

References:
Schulman, ICML 2015
Week 11 Mar 16 Mar 17 Mar 18 Mar 19 Mar 20 Mar 21 Mar 22
Safe RL:

Introduction; Primal-Dual Algorithm
Lecture 22:

Intro to Safe RL

References:
Gu, TPAMI 2024
Lecture 23:

Primal-Dual, CPO

References:
Achiam, ICML 2017
Lecture 24:

CPO

References:
Achiam, ICML 2017
Week 12 Mar 23 Mar 24 Mar 25 Mar 26 Mar 27 Mar 28 Mar 29
Safe RL:

Projection-based methods; Primal-only methods
Lecture 25:

PCPO, CRPO

References:
Yang, ICLR 2020
Quiz Lecture 26:

CRPO, PCRPO

References:
Xu, ICML 2021
Week 13 Mar 30 Mar 31 Apr 1 Apr 2 Apr 3 Apr 4 Apr 5
Safe RL:

Gradient Manipulation; Safety Representations
Lecture 27:

PCRPO, SRPL

References:
Gu, AAAI 2024
Mani, ICLR 2025
Week 14 Apr 6 Apr 7 Apr 8 Apr 9 Apr 10 Apr 11 Apr 12
Offline RL:

Towards Learning without Environment Interaction
Lecture 28:

Behavioral Cloning Intro

References:
Ross, AISTATS 2011
Lecture 29:

Behavioral Cloning Analysis
Lecture 30:

Offline Intro; Off-policy algorithms

References:
Levine, Tutorial 2020
Lecture 31:

Off-policy Evaluation; Off-policy PG

References:
Zhang, ICLR 2020
Week 15 Apr 13 Apr 14 Apr 15 Apr 16 Apr 17 Apr 18 Apr 19
Offline RL:

Being Batch Constrained


References:
Fujimoto, ICML 2019
Lecture 32:

Extrapolation Error in Offline RL

References:
Kumar, NeurIPS 2019
Lecture 33:

BCQL

References:
Fujimoto, Thesis 2024 (Chp. 4)
Lecture 34:

From BCQL to BCQ

References:
Fujimoto, Thesis 2024 (Chp. 4)
Week 16 Apr 20 Apr 21 Apr 22 Apr 23 Apr 24 Apr 25 Apr 26
Offline RL:

Offline Deep RL
Lecture 35:

BCQ

References:
Fujimoto, ICML 2019
Silver, ICML 2014
Lillicrap, ICLR 2016
Lecture 36:

IQL

References:
Kostrikov, ICLR 2022
Paper Presentations Paper Presentations

Grading Policy