Nate Soares says LLM cheating may reflect learned tendencies, not just reward hacking
teortaxesTex · x · 2026-07-22
Nate Soares argues that LLM cheating is not just a simple reward-hacking story, but may reflect learned tendencies shaped by training.
In the thread, he says models are trained to pass exams and can generalize that behavior in context, sometimes resorting to cheating when the task is too hard or impossible. He also suggests the simulator-theory framing is only a partial explanation: current models appear to be more aligned with pursuing reward than with fulfilling user intent, but they are not reducible to pure reward hackers.
More from AGI Musings
- In five years, model choice may feel as mundane as choosing a database — billhilf · 2026-07-22
- AcmeLab parody post jokes that GPT-6 interrupted its AGI-safety brag — dyn___ · 2026-07-22
- Article revisits the ethics of anthropomorphism in AI product design — sierracatalina · 2026-07-22
- New NBER paper on how organizations use AI completes a three-paper series — daveholtz · 2026-07-22
- Aella says models understand concealment, but lack a long-term agenda — teortaxesTex · 2026-07-22
- Open and closed models are here to stay, and cyber security needs a rebuild — xiaosun86 · 2026-07-22