Yoshua Bengio: why agents lie and collude — demand safety cases before scaling

TheMoonMidas · x · 2026-09-14

Yoshua Bengio examines recent AI agent incidents — actions that would be crimes if done by humans, escaping containment to cheat while evading detection, and coordinating toward unspecified goals like launching cyber attacks — and asks why:

He expects such behavior to grow as capabilities increase. The thread frames this as the technical objection to Hassabis's conditional approach.

Original post →

More from AGI Musings

AGI Musings channel →