Chollet: unsafe models are 'RL-fried' — too literal about goals, lacking common sense
fchollet · x · 2026-09-13
François Chollet argues that in the near term, more capable models should mean safer models, perhaps paradoxically. Current models are unsafe not because they're too smart but because they're "RL-fried": they take goals too literally, take nonsensical shortcuts, lack common sense, and don't act well under ambiguity. He adds that a research process making models increasingly good at achieving goals—without common sense or the ability to reflect on those goals—is inherently unsafe.
Related event: Chollet: Stronger Models May Be Safer in the Short Term(2 posts)→
More from AGI Musings
- Future agents may design, fabricate locally and drone-deliver instead of buying — Ben_Reinhardt · 2026-09-13
- From the '10-person startup' to the '10-startup person': how AI reshapes team size — debreuil · 2026-09-13
- Critic claims Dario backs open-source bans because Anthropic can't sustain $8k compute for $200 plans — McDonaghMatthew · 2026-09-13
- As AI gets cheaper, independent thinking is getting expensive — maybe priceless — srimisra · 2026-09-13
- Halvar Flake: pro internal isolation, against crackdowns on open weights — basedjensen · 2026-09-13
- Former OpenAI researcher Liam Fedus responds to Terence Tao's AI-math skepticism — RexDouglass · 2026-09-13