MIT talk explains LLMs from first principles, no transformers needed

vishalmisra · x · 2026-09-27

Vishal Misra released his MIT talk on LLMs, taking a first-principles view of how and why they work without discussing attention or transformer architecture.

His core argument: SFT, RLHF and RL all reshape the underlying distribution, but beneath it all an LLM is still just next-token prediction from that distribution. Understanding this clarifies where model capabilities come from — and their limits.

Related event: MIT Talk Explains LLMs from First Principles, No Transformer Needed(2 posts)→

Original post →

More from Research

Research channel →