Pain direction' interpretability claim doesn't replicate across 32B vs 72B models
rgblong · x · 2026-10-08
- Interpretability researcher rgblong corrects claims about a model's internal 'pain direction': the axis causing 'relieve pain' less than a random direction holds only for the 32B model (56% vs 81%), while at 72B the two are roughly equal (76% vs 74%).
- He concedes the axis is some direction whose behavioral effects differ from matched sadness/fear/random directions, but that doesn't make it a pain direction — and its many anomalous behaviors deepen doubts about how it was defined.
- Overall verdict: 'who knows what this vector or this task means to the model'.
More from Research
- Evolvent AI releases RSIGym and RSI-Index: benchmarking AI self-improvement, Opus 5 leads at 0.4809 — cihangxie · 2026-10-08
- Full Slides Released for ECCV'26 Tutorial on Diffusion Model Post-Training and Alignment — CSProfKGD · 2026-10-08
- Norvig's Classic Essay on Chomsky and the Two Cultures of Statistical Learning Still Reads Fresh in the LLM Era — 3scorciav · 2026-10-08
- DatologyAI open-sources Zephon, cutting data-order noise from 0.82 to 0.05 points when GPU count changes — lmoroney · 2026-10-08
- ProactiveCoach: Hierarchical Guidance Boosts Proactive AI Assistants by 57.1 Points — skku · 2026-10-08
- STEPQuant: 6-bit quantization of Delta-rule recurrent states cuts serving memory by up to 68.7% — zju-community · 2026-10-08