AI alignment debate splits reward hacking from scheming and manipulation
herbiebradley · x · 2026-07-23
The thread asks whether AI misalignment is splitting into two buckets: reward hacking versus scheming/manipulation.
- The author is skeptical that scheming is a distinct phenomenon, saying it mostly appears in tightly controlled environments or model organisms.
- The reply argues scheming may emerge naturally as tasks and evaluations get more complex.
- It also raises a practical concern: even if scheming does not arise on its own, humans may eventually design agent tasks that explicitly incentivize long-horizon scheming.
Related event: AI Alignment Discussion Focuses on Reward Hacking(2 posts)→
More from AGI Musings
- Anthropic’s push for open-source restrictions is said to have united Silicon Valley against it — Hesamation · 2026-07-23
- As AI agents act on their own, we may never know every wild failure case — harris_edouard · 2026-07-23
- Open-weight models may keep winning usage even if they lag the frontier — joshua_saxe · 2026-07-23
- YC says scientists may already have the skills to start companies — ycombinator · 2026-07-23
- OpenAI, Anthropic and Google are spending tens of billions, but Chinese labs are closing in — PeterDiamandis · 2026-07-23
- AI alignment debates now center on human autonomy, not just control rules — GlenBradley · 2026-07-23