Course Description and Objectives
In this course, we will develop the knowledge and skills to understand, implement and apply recent deep reinforcement learning techniques. This course is intended for graduate and advanced undergraduate students.
Specifically, we set the following course objectives:
- Learn about techniques to design and analyse RL algorithms
- Learn about the foundational principles behind recent advances in deep RL
- Gain awareness of the frontier research areas in deep RL
Prerequisites:
- Theory: Probability and statistics; Algorithms; Exposure to AI and/or ML will be helpful
- Implemenation: Fluency in a high level programming language; Familiarity with installing and working with Github packages; Experience with using high-performance compute services will be helpful
Logistics
Time and Location:
- Slot: C
- Class Timings: Tue, Wed, Fri 8 am
- Venue: LHC 518
- Office hours
- Timing: By appointment
- Venue: 513-D, Academic Complex East
Communication:
We will use piazza as the forum for students to ask questions about both the course material as well as logistics. Joining link to be shared.
Teaching Team:
- Instructor: Raunak Bhattacharyya, Assistant Professor, Yardi School of Artificial Intelligence
- Graduate student instructors: Mahesh Keswani, Mohit Chhabra
Course Contents
Below is a tentative list of topics:
- Analysis of tabular RL algorithms
- Approximations
- Policy Search
- Probabilistic Graphical models and RL
- Safe RL
- Offline RL
Textbooks
- Introduction to Reinforcement Learning, Sutton and Barto
- Algorithms for Optimisation, Kochenderfer and Wheeler
- Algorithms for Decision Making, Kochenderfer, Wheeler and Wray
- Numerical Optimisation, Nocedal and Wright
- Neuro-Dynamic Programming, Bertsekas and Tsitsiklis
- Introduction to Markov Decision Processes (under preparation), Puterman
Announcements
Announcements will appear here.Schedule
The weekly schedule will be updated with lecture slides and pointers to additional references as we progress through the course.
| Monday | Tuesday | Wednesday | Thursday | Friday | Saturday | Sunday | |
|---|---|---|---|---|---|---|---|
| Week 1 | Jan 5 | Jan 6 | Jan 7 | Jan 8 | Jan 9 | Jan 10 | Jan 11 |
|
Course overview, Bellman Equations
Reference: Puterman Chapter 5 |
Lecture 1: Course Overview Slides |
Lecture 2: Bellman Equations Slides References: S&B |
Lecture 3: Iterative Policy Evaluation Slides |
Lecture 4: Contraction Proof, Bellman Optimality Intro | |||
| Week 2 | Jan 12 | Jan 13 | Jan 14 | Jan 15 | Jan 16 | Jan 17 | Jan 18 |
|
Contraction Mapping
Reference: Puterman Chapter 5 |
Lecture 5: Bellman Optimality Eqn Proof |
||||||
| Week 3 | Jan 19 | Jan 20 | Jan 21 | Jan 22 | Jan 23 | Jan 24 | Jan 25 |
|
Value Iteration Analysis
Reference: Puterman Chapter 5 |
Lecture 6: Value Iteration |
Lecture 7: Bellman Optimality Contraction |
Lecture 8: Quality of Policy in VI, PI Intro |
||||
| Week 4 | Jan 26 | Jan 27 | Jan 28 | Jan 29 | Jan 30 | Jan 31 | Feb 1 |
|
Policy Iteration, Approximate PI
References: Bertsekas, 2011 Scherrer |
Lecture 9: Policy Iteration Analysis |
Lecture 10: Approximate Policy Iteration Intro |
Lecture 11: API: Dataset Generation |
||||
| Week 5 | Feb 2 | Feb 3 | Feb 4 | Feb 5 | Feb 6 | Feb 7 | Feb 8 |
| Searching in Policy Space |
Lecture 12: API: No Monotonicity References: Pirotta, 2013 Wagner, 2011 |
Lecture 13: Expected Advantage Objective References: Kakade, 2002 |
Lecture 14: Performance Difference Lemma |
||||
| Week 6 | Feb 9 | Feb 10 | Feb 11 | Feb 12 | Feb 13 | Feb 14 | Feb 15 |
| Performance Difference Lemma |
Lecture 15: PDL proof |
Quiz |
Lecture 16: CPI Analysis References: Kakade, ICML 2002 |
||||
| Week 7 | Feb 16 | Feb 17 | Feb 18 | Feb 19 | Feb 20 | Feb 21 | Feb 22 |
|
Conservative Policy Iteration
References: Kakade, ICML 2002 |
Lecture 17: CPI monotonicity |
Lecture18: CPI from PDL | |||||
| Week 8 | Feb 23 | Feb 24 | Feb 25 | Feb 26 | Feb 27 | Feb 28 | Mar 1 |
| Week 9 | Mar 2 | Mar 3 | Mar 4 | Mar 5 | Mar 6 | Mar 7 | Mar 8 |
| Week 10 | Mar 9 | Mar 10 | Mar 11 | Mar 12 | Mar 13 | Mar 14 | Mar 15 |
|
Policy Optimisation: Trust-region methods |
Lecture 19: Parameterised policies; KL divergence constraint References: Kakade, NeurIPS 2001 |
Lecture 20: Optimisation; Linear Objective Quadratic Constraint |
Lecture 21: TRPO References: Schulman, ICML 2015 |
||||
| Week 11 | Mar 16 | Mar 17 | Mar 18 | Mar 19 | Mar 20 | Mar 21 | Mar 22 |
|
Safe RL: Introduction; Primal-Dual Algorithm |
Lecture 22: Intro to Safe RL References: Gu, TPAMI 2024 |
Lecture 23: Primal-Dual, CPO References: Achiam, ICML 2017 |
Lecture 24: CPO References: Achiam, ICML 2017 |
||||
| Week 12 | Mar 23 | Mar 24 | Mar 25 | Mar 26 | Mar 27 | Mar 28 | Mar 29 |
|
Safe RL: Projection-based methods; Primal-only methods |
Lecture 25: PCPO, CRPO References: Yang, ICLR 2020 |
Quiz |
Lecture 26: CRPO, PCRPO References: Xu, ICML 2021 |
||||
| Week 13 | Mar 30 | Mar 31 | Apr 1 | Apr 2 | Apr 3 | Apr 4 | Apr 5 |
|
Safe RL: Gradient Manipulation; Safety Representations |
Lecture 27: PCRPO, SRPL References: Gu, AAAI 2024 Mani, ICLR 2025 | ||||||
| Week 14 | Apr 6 | Apr 7 | Apr 8 | Apr 9 | Apr 10 | Apr 11 | Apr 12 |
|
Offline RL: Towards Learning without Environment Interaction |
Lecture 28: Behavioral Cloning Intro References: Ross, AISTATS 2011 |
Lecture 29: Behavioral Cloning Analysis |
Lecture 30: Offline Intro; Off-policy algorithms References: Levine, Tutorial 2020 |
Lecture 31: Off-policy Evaluation; Off-policy PG References: Zhang, ICLR 2020 |
|||
| Week 15 | Apr 13 | Apr 14 | Apr 15 | Apr 16 | Apr 17 | Apr 18 | Apr 19 |
|
Offline RL: Being Batch Constrained References: Fujimoto, ICML 2019 |
Lecture 32: Extrapolation Error in Offline RL References: Kumar, NeurIPS 2019 |
Lecture 33: BCQL References: Fujimoto, Thesis 2024 (Chp. 4) |
Lecture 34: From BCQL to BCQ References: Fujimoto, Thesis 2024 (Chp. 4) |
||||
| Week 16 | Apr 20 | Apr 21 | Apr 22 | Apr 23 | Apr 24 | Apr 25 | Apr 26 |
|
Offline RL: Offline Deep RL |
Lecture 35: BCQ References: Fujimoto, ICML 2019 Silver, ICML 2014 Lillicrap, ICLR 2016 |
Lecture 36: IQL References: Kostrikov, ICLR 2022 | Paper Presentations | Paper Presentations |