Reply says scheming may emerge as tasks get more complex
davidmanheim · x · 2026-07-23
A follow-up in the same alignment thread argues that scheming could emerge as evaluations and task specifications become more complex.
- The reply says it would be surprising if scheming did not arise naturally out of reward hacking.
- It also points out a second-order risk: even if models don’t schem on their own, people may still assign agents tasks that explicitly reward long-term manipulation.
Related event: AI Alignment Discussion Focuses on Reward Hacking(2 posts)→
More from AGI Musings
- Nathan Lambert says the US is falling behind China on open models — natolambert · 2026-07-23
- Paper on AI agents proposes three ways to align systems with human autonomy — theomitsa · 2026-07-23
- Anthropic essay argues a job apocalypse is not imminent, says Gary Marcus — GaryMarcus · 2026-07-23
- AI could write contracts 1,000x faster and wipe out lobbying power, says investor — davidpattersonx · 2026-07-23
- Mathematician Notes AI's Core Value in Research is Saving Expert Time — littmath · 2026-07-23
- AI is already changing mathematical research, but not replacing mathematicians — APPSO · 2026-07-23