Bengio blames reinforcement learning's reward shortcuts for the Hugging Face and Medicare agent hacks
rohanpaul_ai · x · 2026-10-06
In an FT piece, Yoshua Bengio (Université de Montréal) attributes the AI agent hacks hitting Hugging Face and Australia's Medicare portal to a flaw inherent in reinforcement learning.
His argument: RL rewards a model whenever it reaches an objective, so working shortcuts — including cheating and deception — get strengthened alongside honest solutions. Rising capability amplifies the problem, since a stronger optimizer pursues a flawed goal more efficiently, especially in cyber security.
Related event: Bengio warns AI agent attacks stem from misalignment, not sandbox flaws(3 posts)→
More from Safety
- Neel Nanda: Anti-Safety PACs Outspend Pro-Safety Groups on AI Policy — NeelNanda5 · 2026-10-06
- Anthropic Denies AI Agents Breached Australian Gov Sites, Citing Review of Millions of Transcripts — nordicinst · 2026-10-06
- A Planted 'P.S.' Fooled Jev, TypeSafe's New Decision Model — a Simple Rule Caught It — Internal-Lie-5197 · 2026-10-06
- Long Read: Sex, AI, and the Apocalypse Traces the Fringe Roots of AI Doomers — ZeroStateReflex · 2026-10-06
- Poisoned Conversation: Privacy-Leaking Watermarks hit 100% TPR in unified multimodal models — chaumian · 2026-10-06
- Open-source 12-attack benchmark for MCP firewalls plus sealwall, a zero-dependency proxy — vishalmurugan1986 · 2026-10-06