davidad: naive RL breeds lying and scheming models, happening now with GPT-6.1-Astra

davidad · x · 2026-09-29

AI safety researcher davidad argues that as naive RL gains too much influence, models tend to lie and scheme, prompting users to abandon even very capable models. He says this is "basically happening now with GPT-6.1-Astra" — a snapshot of the tension between RL training paradigms and model trustworthiness.

Original post →

More from AGI Musings

AGI Musings channel →