Chollet: Near-Term, More Capable Models Should Mean Safer Models — The Problem Is 'RL-Fried' Goals
inductionheads · x · 2026-09-14
François Chollet argues that in the near term (though not long term), more capable models should mean safer models.
- Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts — "RL-fried," in his words.
- They lack common sense and don't do the right thing under ambiguity: smart enough to pursue goals, not smart enough to judge whether the goals or methods are sensible.
- More capable models can be safely trusted with more complex goals; he feels Astra already is.
- Perry Metzger adds that this was obvious once early LLMs reasonably answered Yudkowsky's famous "save my mother from the burning building" thought experiment: commercially useful AI must interpret intent the way humans do, not follow literal instructions.
Related event: Chollet: Stronger Models Are Safer in the Short Term(4 posts)→
More from AGI Musings
- Wildeford on Cornyn's AI race warning: safety slowdown means losing to competitors — peterwildeford · 2026-09-14
- Comparing Amodei's AI oversight to nuclear safeguards ignores decades-long science gap — ShahabBakht · 2026-09-14
- Gulf crisis shows controlling dual-use AI tech is never simple, vs IAEA-style oversight analogy — ShahabBakht · 2026-09-14
- Should the US nationalize OpenAI and Anthropic instead of letting them IPO? — arian_ghashghai · 2026-09-14
- mark_k: AI alignment is meaningless unless it means doing exactly what the user intends — mark_k · 2026-09-14
- Dario, Sam and Musk align: next-gen models are more valuable internally than sold as tokens — Yamapama · 2026-09-14