Course Description and Objectives
In this course, we will develop the knowledge and skills to understand, implement and apply recent deep reinforcement learning techniques. This course is intended for graduate and advanced undergraduate students.
Specifically, we set the following course objectives:
- Formulate sequential decision making problems as Markov Decision Processes and map them to suitable RL algorithms.
- Use deep RL algorithms on standard benchmark environments.
- Deploy RL algorithms in various domains in conjunction with other ML techniques.
Prerequisites:
- Theory: Probability and statistics; Algorithms; Exposure to AI and/or ML will be helpful
- Implementation: Fluency in a high level programming language; Familiarity with installing and working with Github packages; Experience with using high-performance compute services will be helpful
Logistics
Time and Location:
- Slot: AC
- Class Timings: Tue/Fri 2-3:20 pm
- Venue: LH313.5
- Office hours
- Timing: By appointment
- Venue: 513-D, Academic Complex East
Communication:
We will use piazza as the forum for students to ask questions about both the course material as well as logistics.
Teaching Team:
- Instructor: Raunak Bhattacharyya, Assistant Professor, Yardi School of Artificial Intelligence
- Graduate student instructors: Mohit Chhabra, Mahesh Keswani, Hasan Mustafa
Course Contents
Below is a tentative list of topics:
- Markov Decision Processes
- Dynamic Programming; Approximations
- Monte Carlo Prediction and Control
- Temporal-Difference Methods
- Exploration-Exploitation Tradeoff
- Policy Gradients
- Imitation learning
- Offline RL
Textbooks
- Introduction to Reinforcement Learning, Sutton and Barto
- Algorithms for Optimisation, Kochenderfer and Wheeler
- Algorithms for Decision Making, Kochenderfer, Wheeler and Wray
- Planning with Markov Decision Processes, Mausam and Kolobov
Announcements
Announcements will appear here.Schedule
The weekly schedule will be updated with lecture slides and pointers to additional references as we progress through the course.
| Monday | Tuesday | Wednesday | Thursday | Friday | Saturday | Sunday | |
|---|---|---|---|---|---|---|---|
| Week 1 | July 21 | July 22 | July 23 | July 24 | July 25 | July 26 | July 27 |
| Course overview |
Lecture 1: Course overview Slides |
||||||
| Week 2 | July 28 | July 29 | July 30 | July 31 | Aug 1 | Aug 2 | Aug 3 |
| Stochastic Sequential Models |
Lecture 2: Hidden Markov Models Slides References: Jurafsky |
Lecture 3: Markov Decision Processes Slides References: S&B (chp 3) Additional Readings: Bellman,1957 |
|||||
| Week 3 | Aug 4 | Aug 5 | Aug 6 | Aug 7 | Aug 8 | Aug 9 | Aug 10 |
| MDPs and Value Functions |
Lecture 4: Value Functions Slides References: S&B (chp 3) |
Lecture 5: MDP Formalisation Slides References: M&K (chp 2) |
|||||
| Week 4 | Aug 11 | Aug 12 | Aug 13 | Aug 14 | Aug 15 | Aug 16 | Aug 17 |
| Policy Evaluation |
Lecture 6: Dynamic Programming Slides References: S&B (chp 4) Additional Readings: Minsky,1961 |
Lecture 7: Monte Carlo Slides References: S&B (chp 5) Additional Readings: Barto & Duff,1994 Michie & Chambers,1968 Widrow, Gupta & Maitra,1973 |
|||||
| Week 5 | Aug 18 | Aug 19 | Aug 20 | Aug 21 | Aug 22 | Aug 23 | Aug 24 |
| Policy Improvement |
Lecture 8: Policy Improvement Slides References: S&B (chp 4) Additional Readings: Puterman & Shin,1978 |
Lecture 9: Policy Iteration and Value Iteration Slides References: S&B (chp 4) Additional Readings: Connell,1989 |
|||||
| Week 6 | Aug 25 | Aug 26 | Aug 27 | Aug 28 | Aug 29 | Aug 30 | Aug 31 |
| Model-Free Prediction and Control |
Lecture 10: Temporal Difference Prediction Slides References: S&B (chp 5&6) Additional Readings: Dayan,1992 Jaakkola, Jordan & Singh,1993 Sutton,1988 Tsitsiklis,1994 |
Lecture 11: Batch TD and MC; Monte Carlo Control Slides References: S&B (chp 5&6) Additional Readings: Singh & Sutton,1996 |
|||||
| Week 7 | Sep 1 | Sep 2 | Sep 3 | Sep 4 | Sep 5 | Sep 6 | Sep 7 |
| Monte Carlo Control |
Lecture 12: Monte Carlo Control Slides References: S&B (chp 5) Additional Readings: Singh & Sutton,1996 Michie & Chambers,1968 |
||||||
| Week 8 | Sep 8 | Sep 9 | Sep 10 | Sep 11 | Sep 12 | Sep 13 | Sep 14 |
| TD Control |
Lecture 13: TD Control Slides References: S&B (chp 6) |
||||||
| Week 9 | Sep 15 | Sep 16 | Sep 17 | Sep 18 | Sep 19 | Sep 20 | Sep 21 |
| Week 10 | Sep 22 | Sep 23 | Sep 24 | Sep 25 | Sep 26 | Sep 27 | Sep 28 |
| Exploration |
Lecture 14: Exploration Slides References: S&B (chp 2) |
||||||
| Week 11 | Sep 29 | Sep 30 | Oct 1 | Oct 2 | Oct 3 | Oct 4 | Oct 5 |
| Week 12 | Oct 6 | Oct 7 | Oct 8 | Oct 9 | Oct 10 | Oct 11 | Oct 12 |
| Approximating Value Functions |
Lecture: HPC Intro, PyTorch Tutorial Slides |
Lecture 15: Approximating value functions Slides References: S&B (chp 9) |
|||||
| Week 13 | Oct 13 | Oct 14 | Oct 15 | Oct 16 | Oct 17 | Oct 18 | Oct 19 |
| DQN |
Lecture 16: Deep Q Networks Slides References: Mnih, 2013 |
Lecture 17: Double Q learning; PG Intro Slides References: Optimizer's Curse van Hasselt, 2010 |
|||||
| Week 14 | Oct 20 | Oct 21 | Oct 22 | Oct 23 | Oct 24 | Oct 25 | Oct 26 |
| Policy Optimisation |
Lecture 18: Policy Optimisation Slides |
||||||
| Week 15 | Oct 27 | Oct 28 | Oct 29 | Oct 30 | Oct 31 | Nov 1 | Nov 2 |
| Policy Gradients |
Lecture 19: Policy Gradients Slides References: S&B (chp 13) |
Lecture 20: Actor-Critic Methods Slides References: A3C, 2015 Additional Readings: Natural Actor-Critic, 2009 |
|||||
| Week 16 | Nov 3 | Nov 4 | Nov 5 | Nov 6 | Nov 7 | Nov 8 | Nov 9 |
| Policy Gradient Extensions |
Lecture 21: Extensions to Policy Search Slides |
Lecture 22: Surrogate Advantage Objective References: TRPO, 2015 PPO, 2017 Additional Readings: NPG |
|||||
| Week 17 | Nov 10 | Nov 11 | Nov 12 | Nov 13 | Nov 14 | Nov 15 | Nov 16 |
| Off-Policy and Max Entropy |
Lecture 23: Model-Based RL; Max-Entropy Slides References: S&B (chp 8) |
Lecture 24: Soft VI; Soft PI; SAC Slides References: Soft Q Learning, 2017 Soft Actor Critic, 2018 |
Grading Policy
Grading policy (tentative, upto ±15% on each component):
- Programming Assignments: 60%
- Exams: 40%
- Audit policy: Only based on exams: To get an audit pass, you should obtain at least 50% absolute score on each and every exam.