COLM 2026 Actionable Interpretability Workshop Accepts 71 Papers, Posters Online
ChrisGPotts · x · 2026-10-09
The Actionable Interpretability Workshop at COLM 2026 accepted 71 papers across two poster sessions (4 as contributed talks), with posters available online. Topics include non-identifiability of steering vectors, guardrail vectors and framing-induced safety failures, fuzzing LLMs for hidden behaviors, localizing attention heads behind LLaVA object hallucinations, and interpretability of latent reasoning models.
More from Research
- SubDGuide: modeler-inspired agentic workflow for mesh-to-SubD reconstruction — ssh4net · 2026-10-10
- Simons Institute launches Fall '27 semester program on diffusions and flows, postdoc fellowships open — giannis_daras · 2026-10-10
- MIT paper: AGI's binding constraint is human verification bandwidth, not intelligence — Afinetheorem · 2026-10-10
- Ex-OpenAI Operator Researcher Previews Real-World RL Gains on High-Precision Tasks — ZhaoMandi · 2026-10-10
- Stanford's Mulligan targets failure states for on-robot learning, beats uniform collection across 2,550 blind evals — ZhaoMandi · 2026-10-10
- RCT: AI as tutor raises test scores that persist; letting AI write for you fades within a week — emollick · 2026-10-10