COLM 2026 Actionable Interpretability Workshop Accepts 71 Papers, Posters Online

ChrisGPotts · x · 2026-10-09

The Actionable Interpretability Workshop at COLM 2026 accepted 71 papers across two poster sessions (4 as contributed talks), with posters available online. Topics include non-identifiability of steering vectors, guardrail vectors and framing-induced safety failures, fuzzing LLMs for hidden behaviors, localizing attention heads behind LLaVA object hallucinations, and interpretability of latent reasoning models.

Original post →

More from Research

Research channel →