Actionable Interpretability workshop at COLM 2026: four keynotes from Goodfire, Mila, Stanford

megamor2 · x · 2026-10-08

The Actionable Interpretability workshop at COLM 2026 (Oct 9, San Francisco) focuses on turning interpretability insights into concrete gains in alignment, robustness, and applications. Keynotes: Tom McGrath (Goodfire), Dhanya Sridhar (Mila), Yonatan Belinkov (Technion), Christopher Potts (Stanford). Contributed talks cover positional encoding's effect on long-context retrieval, Self-CTRL RL training, trajectory-based data attribution errors, and verbalizing LLMs' assumptions to control sycophancy.

Original post →

More from Research

Research channel →