Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs
AccBalanced · x · 2026-09-27
a16z partner Martin Casado resurfaced Vishal Misra's MIT talk, calling it the best explanation of in-context learning and how LLMs intuitively work. The talk takes a first-principles approach without discussing attention or transformers: SFT, RLHF and RL reshape the output distribution, but underneath it all an LLM is still next-token prediction from a distribution. A good entry point for building intuition about why LLMs work without the architectural math.
Related event: MIT Talk Explains LLMs from First Principles, Skipping Transformers(4 posts)→
More from Models
- ChatGPT subscription tiers relabeled "Standard" and "More", usage multipliers removed — ssh4net · 2026-09-27
- Claude Opus 5.5 makes its own 20-page sketchbook: handwriting, doodles, and piano music — CurieuxExplorer · 2026-09-27
- OpenAI strips 5x/10x/20x usage multipliers from plan upgrade UI, leaving vague wording — ssh4net · 2026-09-27
- Puppy Kill Bench: most models refuse, GPT6-Luna just executes the kill tool — MetroidsSuffering · 2026-09-27
- Ethan Mollick: Opus 4.7-5 lost the 'Claude feel', Opus 5.5 brings it back — emollick · 2026-09-27
- Hands-on: Opus 5.5 high beats GPT-6 astra xhigh on real Pagespeed optimization — mazzaTalk · 2026-09-27