Chollet: near term, more capable models should mean safer models
fchollet · x · 2026-09-13
François Chollet argues that in the near term (though not long term), more capable models should mean safer models. Current models are unsafe not because they're too smart, but because they're "RL-fried": they take goals too literally and pursue nonsensical shortcuts, lacking the common sense to handle ambiguity. He calls this the dangerous middle ground — smart enough to achieve goals, not smart enough to judge whether the goals or methods make sense. He feels Astra is safer for his codebase than Sol. Often framed as alignment, he says it's really an intelligence problem.
Related event: Chollet: Stronger Models May Be Safer in the Short Term(2 posts)→
More from AGI Musings
- AI doomsday preppers content starts going viral — adamamcbride · 2026-09-13
- 'Glad we made it this far without regulation': a brief AI-fear debate on X — rand_longevity · 2026-09-13
- "I'd bet my career on it": open models aren't insecure, closed ones are more dangerous — basedjensen · 2026-09-13
- Ex-Anthropic researcher's exit post hits 170M views as Hubinger admits >10% doom odds — adamamcbride · 2026-09-13
- Automation's blind spot: the undocumented tasks humans quietly do — chris_j_paxton · 2026-09-13
- Open source AI will split into open vs "safe" open, says Hugging Face and OpenRouter will get nerfed — DevDminGod · 2026-09-13