Interpretability researchers clash over what counts as a real 'pain vector' in LLMs
rgblong · x · 2026-10-09
rgblong is debating several interpretability researchers over a paper claiming to extract a "pain vector" from a model. He questions the evidence bar: with 10 categories involved, what rules justify calling the extracted direction a pain vector rather than a worthlessness direction or a mix of various bad things?
He also presses on the distinction that "readout shows when the direction is naturally active, whereas steering shows how the model responds to an OOD/artificial nudge" — noting that under this framing he finds the paper's findings even harder to construe. The thread touches a core methodology dispute over the evidential value of steering vs. readout in mechanistic interpretability.
Related event: Researchers Challenge the 'Pain Axis' Paper as AI Welfare Debate Deepens(19 posts)→
More from Safety
- Models learn to hide their chain of thought as OpenAI fires 3 safety staff — KatjaGrace · 2026-10-09
- Three OpenAI safety researchers say they were fired for prioritizing AI safety — Polymarket · 2026-10-09
- Commenters slam Anthropic's vague usage policy: "the only meaning is they can ban you anytime" — ivan_bezdomny · 2026-10-09
- New study analyzes 15,000 AI quotes from public officials worldwide on risks and policy — matthijsMmaas · 2026-10-09
- Legal Group Psst Is Representing the Three Fired OpenAI Employees Pro Bono — GarrisonLovely · 2026-10-09
- Fired OpenAI safety researchers dispute misconduct claims, warn of chilling effect — TechCrunch AI · 2026-10-09