Anil Seth: LLM 'Pain Vectors' Can Aid Interpretability and AI Safety

anilkseth · x · 2026-09-21

Part 2 of Anil Seth's thread on the Camhberg et al. preprint: he praises the identification of functional 'pain' vectors in LLMs for interpretability and AI safety, especially experiments where models 'pay a cost' (worse answers, deleting user files) to obtain 'relief'. See the thread's first post for the full take.

Related event: Pain-vector LLM paper authors and Anil Seth push back on misreadings(13 posts)→

Original post →

More from AGI Musings

AGI Musings channel →