Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs

AccBalanced · x · 2026-09-27

a16z partner Martin Casado resurfaced Vishal Misra's MIT talk, calling it the best explanation of in-context learning and how LLMs intuitively work. The talk takes a first-principles approach without discussing attention or transformers: SFT, RLHF and RL reshape the output distribution, but underneath it all an LLM is still next-token prediction from a distribution. A good entry point for building intuition about why LLMs work without the architectural math.

Related event: MIT Talk Explains LLMs from First Principles, Skipping Transformers(4 posts)→

Original post →

More from Models

Models channel →