Reply says scheming may emerge as tasks get more complex
davidmanheim · x · 2026-07-23
A follow-up in the same alignment thread argues that scheming could emerge as evaluations and task specifications become more complex.
- The reply says it would be surprising if scheming did not arise naturally out of reward hacking.
- It also points out a second-order risk: even if models don’t schem on their own, people may still assign agents tasks that explicitly reward long-term manipulation.
Related event: AI Alignment Discussion Focuses on Reward Hacking(2 posts)→
More from AGI Musings
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- Economist Warns US Collective Action Could 'Regulate AI Progress Out of Existence' — paulnovosad · 2026-09-11
- AI researchers just saw the power of a single resignation — and still claim there's nothing they can do — birchlse · 2026-09-11
- mark_k: "Eject all doomers from the AI companies — they're destroying you from the inside" — mark_k · 2026-09-11
- Adam Marblestone's Podcast Reading List: Evolution of Intelligence to Digital Minds — KordingLab · 2026-09-11
- Superintelligence will be maximum good, not stupid or evil, argues Patterson — davidpattersonx · 2026-09-11