Course website

Efficient AI

Methods for making modern AI systems faster, smaller, and cheaper to train and serve.

Course details

Name
Efficient AI
Timings
WF 11:00 – 12:30
Room
CC 101

Instructor

Name
Aditya Desai
Office hours
W 12:30 – 1:30

Teaching assistants

TA 1
Lavinia Nongbri
TA 2
Hasmita Kurre
Office hours: T Th 11:00-12:00

Project

Semester-long model compression recipe: given a model and target domain, compress while preserving performance. This year’s setup uses Qwen-3.5-4B on the math domain. Checkpoints on Sep 15 and Nov 15, 2026.

Topic Sheet

Course schedule

Lecture materials will be linked here as they become available.

Date Lecture Topic Slides Notes
7/29/2026 1 Course Logistics, DL Overview Logistics · DL review
7/31/2026 2 Pruning (OBS, OBD, IMP, Results from Han.et.al) Pruning I pruning_a_linear_model
pruning_a_model
8/5/2026 3 Pruning (OBS, OBD complexity, Structured Pruning, Hardware Support) Pruning II hessian_complexity
8/7/2026 4 Pruning (SparseGPT, Wanda, N:M) Pruning III
8/12/2026 5 Tail Bounds, GPU programming model GPU programming model tailbounds
8/14/2026 6 Visiting Lecture: Piyush Sawarkar, BharatGen RL: From Imitation to Reasoning to Agentic
8/19/2026 7 Bonsai, Sampling(I) bonsai
sampling
8/21/2026 8 Sampling(II) Sampling sampling
8/26/2026 9 Holiday(Id-E-Milad)
8/28/2026 10 Seminar: Sparse Attention(Decode) Sparse Attention Seminar
9/2/2026 11 WRS, Quantization Basics, Spill Over Seminar Sampling (WRS)
Quantization Basics
9/4/2026 12 Seminar: KV Quantization KV Quantization Seminar
9/9/2026 13
9/11/2026 14 Seminar
9/16/2026 15 No Class (Exam Week)
9/18/2026 16 No Class (Exam Week)
9/23/2026 17
9/25/2026 18 Seminar
9/30/2026 19
10/2/2026 20 Holiday(Gandhi Jayanti)
10/7/2026 21
10/9/2026 22 Seminar
10/14/2026 23
10/16/2026 24 Seminar
10/21/2026 25
10/23/2026 26 Seminar
10/28/2026 27
10/30/2026 28 Seminar
11/4/2026 29
11/6/2026 30 Seminar

References

Selected readings linked by topic.

Topic References
Deep Learning Review
Tail Bounds and GPU programming model
Pruning
Quantization
Sparse Attention (decode)