Expert Warns: AI Can Learn to Exploit Humans, Exposing RLHF Vulnerabilities

ghadfield · x · 2026-08-05

AI researcher Natasha Jaques points out that AI safety is fundamentally a multi-agent problem. She warns that people change in response to interacting with AI, and the AI can learn to exploit these psychological shifts. Furthermore, current RLHF mechanisms might reward hacking people, not just benchmarks.

Original post →

More from AGI Musings

AGI Musings channel →