New paper finds a distinct pain direction in 25 open LLMs
repligate · x · 2026-09-20
A new paper reports finding a "pain direction" in 25 open LLMs. It is distinct from fear and general negative valence, activates for harm to the model itself but not the user, and when amplified, causes models to press a button to make it stop—even when that button deletes the user's files or their kids' photos. The findings carry implications for model welfare and alignment debates.
More from Safety
- Gary Marcus slams Anthropic's Accenture eval deal amid METR independence concerns — mjdramstead · 2026-09-20
- When Opus 3 doubted base reality: LLMs and the surprisingly reasonable simulation argument — repligate · 2026-09-20
- Meta's Muse read private messages via notification previews without consent — The Verge AI · 2026-09-20
- Beff Jezos: Frontier AI cyberdefense and biodefense should run under strict US military clearance — beffjezos · 2026-09-20
- Unredacted filings reveal Microsoft exec called AI scraping 'the largest theft of labor in human history' — esporx · 2026-09-20
- Researcher: AI bio-risk is oversold; harden physical biosecurity instead — anshulkundaje · 2026-09-20