Stanford launches MS&E 319 on efficient LLMs: MoE, quantization, speculative decoding

aminkarbasi · x · 2026-09-21

Amin Karbasi and colleagues (@AnayMehrotra, gvelegkas, Amin Saberi) are teaching Stanford's MS&E 319: Efficient Generative Language Models this fall, tackling how to pick training objectives, architectures and inference algorithms under limited compute budgets.

Topics include efficient pre-training, mixture-of-experts, attention and KV-cache compression, quantization, speculative decoding, LoRA, RLHF, DPO and distillation. Lecture materials and recordings will be posted as the course unfolds — 'just buy more GPUs' won't be the only answer.

Original post →

More from Companies & People

Companies & People channel →