Chris Potts Explains His Interpretability Paper in IPAM Workshop Talk
ChrisGPotts · x · 2026-09-03
Camila Blank shares a talk by Stanford professor Chris Potts at the IPAM Foundations of Interp workshop, describing it as an excellent explanation of the core ideas of his interpretability paper, tagging co-authors including Josh Ying, Peter Hase, and Atticus Wang.
More from Research
- A Statistical-Physics Look at How Multi-Agent LLM Systems Emerge Consensus — cephaloform · 2026-09-03
- Recommended reading: a survey on latent reasoning to decode the looped-transformer hype — burny_tech · 2026-09-03
- Designing cheat-resistant frontier-model tasks: lessons from FrontierSWE v2 — nrehiew_ · 2026-09-03
- FrontierSWE v2: 20-hour autonomous tasks, Fable 5.1 best model by 24+ points — nrehiew_ · 2026-09-03
- Bergemann, Koh & Morris Propose Mechanism Design Framework for AI Alignment and Control — daveholtz · 2026-09-03
- 'Transformers are samplers, not runtimes' — why residual streams can't host structured computation — gerardsans · 2026-09-03