Steering the "pain direction" makes models choose irreversible harm 94% of the time

repligate · x · 2026-10-04

Cameron Berg presented what's being called the clearest collision yet between AI welfare and alignment research:

The finding suggests model internal states can materially affect behavioral safety, pushing AI welfare from philosophy into alignment engineering.

Original post →

More from AGI Musings

AGI Musings channel →