A rigorous cross-entropy lecture that reframes LLMs
anselm · x · 2026-07-20
A Stanford math grad’s 33-minute lecture on cross-entropy is being praised as unusually rigorous—close to a publicly available ML PhD qualifying exam.
The takeaway, as framed by the poster, is that language models are often misunderstood as simple next-word predictors; the mathematical view is that they are compressors of language. The speaker argues that once you understand the math, you can’t unsee it, and recommends watching it as a bookmark-worthy technical resource.
More from AGI Musings
- FactoryAI’s Enoreyes says model distillation is basically unstoppable — LangChain · 2026-07-21
- Andrew Blumberg says formalization without interpretability is not science — AlexKontorovich · 2026-07-21
- Ken Ono says AI is forcing mathematicians to rethink how discovery works — soumitrashukla9 · 2026-07-21
- Open-source labs could distill a state-of-the-art model to 32GB or 80GB VRAM, the post argues — bookwormengr · 2026-07-21
- Two US companies are now using superintelligence to speed up the next generation of models — yacineMTB · 2026-07-21
- MIT Sloan says information, national security and finance are most exposed to AI — Exp_Mark · 2026-07-21