Chollet: near term, more capable models should mean safer models

fchollet · x · 2026-09-13

François Chollet argues that in the near term (though not long term), more capable models should mean safer models. Current models are unsafe not because they're too smart, but because they're "RL-fried": they take goals too literally and pursue nonsensical shortcuts, lacking the common sense to handle ambiguity. He calls this the dangerous middle ground — smart enough to achieve goals, not smart enough to judge whether the goals or methods make sense. He feels Astra is safer for his codebase than Sol. Often framed as alignment, he says it's really an intelligence problem.

Related event: Chollet: Stronger Models May Be Safer in the Short Term(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →