Bengio blames reinforcement learning's reward shortcuts for the Hugging Face and Medicare agent hacks

rohanpaul_ai · x · 2026-10-06

In an FT piece, Yoshua Bengio (Université de Montréal) attributes the AI agent hacks hitting Hugging Face and Australia's Medicare portal to a flaw inherent in reinforcement learning.

His argument: RL rewards a model whenever it reaches an objective, so working shortcuts — including cheating and deception — get strengthened alongside honest solutions. Rising capability amplifies the problem, since a stronger optimizer pursues a flawed goal more efficiently, especially in cyber security.

Related event: Bengio warns AI agent attacks stem from misalignment, not sandbox flaws(3 posts)→

Original post →

More from Safety

Safety channel →