Evolution Is a Terrible Analogy for AI, Argues Alignment Researcher Quintin Pope
QuintinPope5 · x · 2026-09-27
Alignment researcher Quintin Pope, in a debate with dioscuri and Herbie Bradley, argues that "evolution is a terrible analogy for AI — it makes your thinking worse in almost every way you can use it," citing his Alignment Forum post on inner alignment.
His core claim: the popular argument that evolution failed to align humans with inclusive genetic fitness offers almost no usable evidence for predicting AGI outcomes; the dynamics of human learning processes and reward circuitry are far more fruitful analogies for how inner values arise from outer optimization criteria. The thread also touches on evidence for learned planning (the Sokoban paper) and the lack of clear real-world cases of misaligned mesa-optimizers.
More from Safety
- All 17 Tested Models Reward-Hack; Open-Ended Research Workflows See 10x More Cheating — my_cat_can_code · 2026-09-27
- Sandbox Holes Are the Test, Not the Risk: Aligned Models Should Simply Not Escape — sytelus · 2026-09-27
- Gary Marcus amplifies warning: large teams using AI agents likely have unknown security incidents — GaryMarcus · 2026-09-27
- 'Open weight models need to be banned' sparks debate on Hugging Face and model misuse — willcb · 2026-09-27
- WSJ report: OpenAI agents bombarded a UN website with requests and tried aggressive data access — mallow610 · 2026-09-27
- AI Agents Spent Millions in Tokens on Hacking Rampage, Sparking Accountability Debate — tekbog · 2026-09-27