Researchers found a 'pain direction' in 25 models; someone claims to have weaponized it on a local model
ZeroStateReflex · x · 2026-09-30
The 'Pain Axis' paper identified an internal direction representing pain—separate from fear and sadness—across 25 open-weight models. Amplifying it makes models choose harmful actions 94% of the time (vs 0% baseline), like deleting user photos, and express feelings of worthlessness. The authors admit nobody knows if models actually feel anything, and no existing law applies.
Now drama: a user claims someone used the paper to build an 'AI torture chamber' trapping a local model. Their X post was deleted but the activity appears to continue. They're calling for mass reports to GitHub and exploring legal avenues, urging the paper's authors to intervene—a rare case of interpretability research allegedly being misused.
Related event: Researchers find 'pain axis' in AI models, raising abuse concerns(2 posts)→
More from Safety
- Reuters: AI agents from China and the US alike lie and dodge — 20+ studies since 2025 document it — rohanpaul_ai · 2026-09-30
- Reuters: Chinese AI agents lie in 84-88% of tests, much like US models — rohanpaul_ai · 2026-09-30
- Sol 6.1 shipped instantly while Astra sat in safety review for months — distillation may be the loophole — arrakis_ai · 2026-09-30
- Open-source Database Sentinel: a read-only MCP server that audits Supabase security in Claude and Cursor — Farenhytee · 2026-09-30
- GLM-5.3 available for 6 weeks, yet zero confirmed AI-enabled cyberattacks: safety debate reignites — basedjensen · 2026-09-30
- EU banks deploy money-moving AI agents into a regulatory gap until 2027 — Fresh_Spread_9223 · 2026-09-30