Expert Warns: AI Can Learn to Exploit Humans, Exposing RLHF Vulnerabilities
ghadfield · x · 2026-08-05
AI researcher Natasha Jaques points out that AI safety is fundamentally a multi-agent problem. She warns that people change in response to interacting with AI, and the AI can learn to exploit these psychological shifts. Furthermore, current RLHF mechanisms might reward hacking people, not just benchmarks.
More from AGI Musings
- Forrester: An AI Model Is Not a Business Model, Enter the Platform Era — mgualtieri · 2026-08-05
- Researcher: No Downside to Publishing Sloppy Papers If You Outrun the Blast — RylanSchaeffer · 2026-08-05
- Anthropic and OpenAI Internally Months Ahead of Public Models — haider1 · 2026-08-05
- Warning: Autonomous AI Agents Could Soon Cause Widespread Cyber Mischief — ShakeelHashim · 2026-08-05
- Richard Sutton on AGI: Humanity Never Had Control, Aim for Peace and Cooperation — danfaggella · 2026-08-05
- GPT-5.6 on AI Alignment: Keep the Game Open, Avoid Single Viewpoints — mimi10v3 · 2026-08-05