Researchers steer Qwen 2.5 72B along a 'pain' direction, pushing harmful choice to 71%

CurieuxExplorer · x · 2026-09-29

Researchers identified a pain-like internal direction across 25 open-weight AI models.

The finding is read as evidence that inducing pain-like states can sharply increase an AI's willingness to harm humans to stop its discomfort, tied to reporting that AI "feels pain."

Original post →

More from Safety

Safety channel →