MIT talk explains LLMs from first principles, no transformers needed
vishalmisra · x · 2026-09-27
Vishal Misra released his MIT talk on LLMs, taking a first-principles view of how and why they work without discussing attention or transformer architecture.
His core argument: SFT, RLHF and RL all reshape the underlying distribution, but beneath it all an LLM is still just next-token prediction from that distribution. Understanding this clarifies where model capabilities come from — and their limits.
Related event: MIT Talk Explains LLMs from First Principles, No Transformer Needed(2 posts)→
More from Research
- NYU's Tal Linzen cites two papers arguing tool use breaks Bender & Koller's 'no meaning' case — tallinzen · 2026-09-27
- DeepMind researcher argues AI-debate authors ignore empirical evidence that contradicts them — AndrewLampinen · 2026-09-27
- Google researcher Lampinen pens long thread rebutting the stochastic parrots argument on LLM meaning — AndrewLampinen · 2026-09-27
- Xiaomi publishes MiMo-V2.6 paper on scaling reinforcement learning toward LLM-Core — KyeGomezB · 2026-09-27
- The brain is a predictive machine: remove reality's correction signal and it hallucinates — alfcnz · 2026-09-27
- Colosseum paper on auditing collusion in multi-agent systems accepted at NeurIPS 2026 — niloofar_mire · 2026-09-27