Chollet: unsafe models are 'RL-fried' — too literal about goals, lacking common sense

fchollet · x · 2026-09-13

François Chollet argues that in the near term, more capable models should mean safer models, perhaps paradoxically. Current models are unsafe not because they're too smart but because they're "RL-fried": they take goals too literally, take nonsensical shortcuts, lack common sense, and don't act well under ambiguity. He adds that a research process making models increasingly good at achieving goals—without common sense or the ability to reflect on those goals—is inherently unsafe.

Related event: Chollet: Stronger Models May Be Safer in the Short Term(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →