Anil Seth: LLM 'Pain Vectors' Can Aid Interpretability and AI Safety
anilkseth · x · 2026-09-21
Part 2 of Anil Seth's thread on the Camhberg et al. preprint: he praises the identification of functional 'pain' vectors in LLMs for interpretability and AI safety, especially experiments where models 'pay a cost' (worse answers, deleting user files) to obtain 'relief'. See the thread's first post for the full take.
Related event: Pain-vector LLM paper authors and Anil Seth push back on misreadings(13 posts)→
More from AGI Musings
- Astro Teller: In 15 years we'll talk about hacking biology, not AI — Scobleizer · 2026-09-21
- Pensions funding the apocalypse? Open models closing the gap put OpenAI/Anthropic premium in question — AlexTensor · 2026-09-21
- sarahookr: Unpredictable pricing and IP concerns swing the pendulum toward custom AI workflows — sarahookr · 2026-09-21
- Naval: AI models trained on open web should be opened after ~12 months — rohanpaul_ai · 2026-09-21
- Jev and the bittersweet lesson: can differently sized models serve different apps? — amankhan · 2026-09-21
- Specialized Models Will Beat General LLMs: The Coming 'PC AI' Era — markjeffrey · 2026-09-21