MIT Talk: Strip Away Attention and LLMs Are Still Next-Token Predictors
soumitrashukla9 · x · 2026-09-27
A video of Vishal Misra's MIT talk offers a first-principles view of how and why LLMs work, without touching attention or transformers: SFT/RLHF/RL reshape the distribution, but underneath it all the model is still next-token prediction. The retweeter's take: the "it's just predicting the next token" crowd wasn't wrong — they just underestimated how powerful this framework is.
More from Research
- NBER paper: rising cognitive skill productivity explains US within-occupation wage inequality — soumitrashukla9 · 2026-09-27
- New RL trick trades variance for bias, stabilizing tiny-batch training runs — willcb · 2026-09-27
- Quail project on building with agents: stand on battle-tested community work — charles_irl · 2026-09-27
- 17M-parameter model beats frontier LLMs 80% of the time after cheap synthetic-data finetuning — max_paperclips · 2026-09-27
- Why the AI Scaling Hypothesis May Never Be Falsified: A Duhem–Quine Argument — burny_tech · 2026-09-27
- NYU's Tal Linzen cites two papers arguing tool use breaks Bender & Koller's 'no meaning' case — tallinzen · 2026-09-27