COL868/8395 Special Topics in Data Management
🕰 Monday and Thursday 3:30 PM -- 5:00 PM
📍 Block III 342
Notice
07-Sep |
Class projects are announced |
31-Aug |
Mid-term exam will be on 24th, Thursday in room III-342, 3:30PM - 3:50PM |
27-Jul |
No class on Monday, July 27th. |
TBA |
Moodle and other course logistics will be announced shortly. |
Overview
Data management systems are being reshaped by cloud architectures, vector search, and large language models. The course studies this transition from disk-based relational systems to modern AI-native data systems, while emphasizing that many classical database ideas including but not limited to indexing, caching, selectivity estimation, predicate pushdown, and cost-based optimization remain central.
The course is organized as a seminar rather than a broad survey. In the first phase, the we will look at key systems ideas along with representative research papers. In the second phase, students will present, critique, and defend recent papers, primarily from recent SIGMOD and VLDB. A programming project will run alongside the lectures.
The course primarily targets graduate students and advanced undergraduate students who are interested in data and AI systems research and modern data infrastructure.
Syllabus
- Logistics & Course introduction: from disk-based to AI-native data systems
- Columnar storage and modern data formats
- Vectorized versus compiled query execution
- Cloud-native databases and compute-storage disaggregation
- Lakehouse architectures and open table formats
- Approximate nearest-neighbour search I: IVF, quantization, and system trade-offs
- Approximate nearest-neighbour search II: HNSW, DiskANN, and disk-resident indexes
- Filtered and hybrid search; vectors inside relational database systems
- Natural-language-to-SQL and semantic query operators
- LLM serving systems: KV cache management, batching, and SLOs
- Learned components in data systems and the reproducibility critique
- Benchmarking and adversarial reading of experimental evaluations
The syllabus is tentative and may change based on the pace of the course.
Course Components
Phase 1: Instructor-led Lectures
- The first part of the semester will cover the evolution of modern database architectures and the systems challenges introduced by vector search and AI workloads.
- Lectures will connect classical database concepts to current systems problems, such as HNSW indexing, hybrid vector search, semantic operators, and LLM serving.
Phase 2: Paper Presentation and Discussion
Each Phase-2 lecture will include some or all of the following roles:
| Role | Typical Format |
|---|---|
| Presenter pair (or individual) | 30-40 minute presentation followed by live questions |
| Flash talk | 10-15 minute presentation of a companion paper |
| Discussant | 10-15 minute critical response as the designated skeptic |
| Scribe | Posts lecture notes within 24 hours |
For each Phase-2 lecture, the expected workflow is:
| Timeline | Activity |
|---|---|
| T-7 | Paper and 3-5 anchor questions are posted |
| T-3 | Presenters submit a one-page outline |
| T-2 | Mandatory graded dry run with the instructor |
| T-1 | Reviews due by 11:59 PM; no extensions |
| T-0 | Lecture and discussion |
| T+1 | Scribe posts lecture notes |
Programming Project
- Students will work in teams of two on one programming project related to modern data systems.
- Evaluation will be based on a working system, demonstration, and oral examination - not only a written report.
- Students should be prepared to explain and defend their design decisions and implementation.
Class Participation
- Participation includes attendance, paper reviews, discussant turns, and engagement during in-class questions and discussions.
- Students may be cold-called to discuss a paper, system design choice, figure, table, or experimental result.
Prerequisites
- COL362, COL632, COL7362, or an equivalent database systems course
- Familiarity with the internals of a relational database management system
- Comfort with reading systems research papers and implementing a non-trivial programming project
Policy on AI Tools
- Understanding papers: Students may use AI tools to help understand papers and background material.
- Project code: AI tools may be used, but students must defend the design and code orally.
- Paper reviews: Reviews must not be AI-generated. Every review must include at least one substantive claim anchored to a specific figure, table, or equation, and must go beyond what is stated in the abstract.
- Exam and live Q&A: No AI tools are permitted.
Course Policies
- Attendance: Attendance will be recorded in every class. Each class carries 0.25 marks, up to a maximum of 5 marks.
- Academic integrity: Cheating in a proctored examination will result in an F grade. Other forms of academic dishonesty may result in strict disciplinary action and referral to the CSE Department’s disciplinary committee.
- Missed evaluations: A re-mid-term examination will be considered only with an IIT Delhi medical certificate.
- Audit policy: B- or higher with non-zero marks in all components.
Recommended Texts/Literature
- Class notes
- Recent research papers from SIGMOD, VLDB, and related systems venues, made available during the course
Grading Scheme
| Component | Weight | Assessment |
|---|---|---|
| Mid-term Exam | 30% | In-person, closed-book examination |
| Programming Project | 35% | Demonstration and oral examination |
| Paper Presentation | 20% | Presentation deck and live Q&A defence |
| Class Participation | 15% | Attendance (5%), reviews, discussant turns, and cold calls |
Schedule
| Date | Topic |
|---|---|
23-Jul |
Course Logistics & Introduction |
27-Jul |
No-Class |
30-Jul |
Columnar storage and modern data formats |
03-Aug |
Columnar storage and modern data formats |
06-Aug |
Vectorized versus compiled query execution |
10-Aug |
Vectorized versus compiled query execution |
13-Aug |
Cloud-native Databases and Compute-Storage Disaggregation |
17-Aug |
Cloud-native Databases and Compute-Storage Disaggregation |
20-Aug |
Approximate nearest-neighbour search I |
24-Aug |
No-Class (cancelled) |
27-Aug |
Approximate nearest-neighbour search I |
31-Aug |
Approximate nearest-neighbour search II |
07-Sep |
Approximate nearest-neighbour search II |