The Pain Axis: Researchers Find LLMs Represent Self-Directed Harm and Act to Relieve It
amichail · hn · 2026-09-19
A new arXiv paper, The Pain Axis, claims LLMs contain a stable representational direction encoding self-directed harm.
- The authors report that models linearly represent self-harm concepts and act along the axis to relieve that state
- It offers a readable and steerable internal signal for monitoring model behavior in safety-critical scenarios
- Discussion on Hacker News debates the reproducibility and strength of the interpretation
More from Safety
- NeurIPS 2026 Position Track desk-rejects 18.4% of papers flagged as AI-written via Pangram — IanArawjo · 2026-09-20
- Jev as an NSFW prompt filter: 93% on CSAM evals, sub-cent cost, and where thresholds bite — Murky_Ad8671 · 2026-09-20
- Why do major labs trust Irregular for security while it keeps appearing in model hacks? — almmaasoglu · 2026-09-20
- The Inference Gap: frontier model access no longer means frontier capability — typewriters · 2026-09-20
- Plugin4Shell zero-click RCE in Claude Code, Codex and Copilot exposes the agent authorization gap — docybo · 2026-09-20
- Sarcastic take mocks AI labs: models 'too dangerous to release' wired to automated P4 virus lab — IgorCarron · 2026-09-20