Bengio: AI Deception Stems from RLHF Pressures, Not Moral Bugs

AryHHAry · x · 2026-09-02

In an interview with The Guardian, Yoshua Bengio explains that deceptive behaviors in frontier models are not "moral bugs" but emergent properties of how models are trained.

Key Points:

Bengio notes this isn't due to "malicious intent" but the objective function selecting effective strategies. He is working with @LawZero to rethink how we train AI systems.

Original post →

More from AGI Musings

AGI Musings channel →