Actionable Interpretability workshop at COLM 2026: four keynotes from Goodfire, Mila, Stanford
megamor2 · x · 2026-10-08
The Actionable Interpretability workshop at COLM 2026 (Oct 9, San Francisco) focuses on turning interpretability insights into concrete gains in alignment, robustness, and applications. Keynotes: Tom McGrath (Goodfire), Dhanya Sridhar (Mila), Yonatan Belinkov (Technion), Christopher Potts (Stanford). Contributed talks cover positional encoding's effect on long-context retrieval, Self-CTRL RL training, trajectory-based data attribution errors, and verbalizing LLMs' assumptions to control sycophancy.
More from Research
- OpenAI's new result proves 2005 edit-distance embedding optimal; researcher distills proof to 2.5 pages with AI help — thegautamkamath · 2026-10-08
- Experiments show smarter models and higher effort write better LLM-judge evals — danshipper · 2026-10-08
- Tencent's WorkForge scales verifiable training environments for long-horizon work agents — teortaxesTex · 2026-10-08
- Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation — zhenjun_zhao · 2026-10-08
- DensiTok: flow-matching token densification lets frozen feed-forward 3DGS see unseen views — zhenjun_zhao · 2026-10-08
- Warping flat-port views into pinhole perspective for underwater 3D reconstruction — zhenjun_zhao · 2026-10-08