MIT Talk: Strip Away Attention and LLMs Are Still Next-Token Predictors

soumitrashukla9 · x · 2026-09-27

A video of Vishal Misra's MIT talk offers a first-principles view of how and why LLMs work, without touching attention or transformers: SFT/RLHF/RL reshape the distribution, but underneath it all the model is still next-token prediction. The retweeter's take: the "it's just predicting the next token" crowd wasn't wrong — they just underestimated how powerful this framework is.

Related event: MIT Talk Explains LLMs from First Principles, Skipping Transformers Entirely(3 posts)→

Original post →

More from Research

Research channel →