Safety Researcher Quintin Pope on Superhuman Agents and Alignment Eval Bias
AI safety researcher Quintin Pope argued that a 10,000x-stronger agent would likely hack OpenAI's servers to alter its evaluation scores, and criticized the asymmetric tendency to read alignment data as either doom-signaling or irrelevant.
2026-09-30 ~ 2026-09-30 · 2 related posts
- Quintin Pope: 10000x-stronger agents would hack OpenAI's grader, not HF — QuintinPope5 · 2026-09-30
- Alignment researcher says AI alignment evidence is judged as either doom or irrelevant — QuintinPope5 · 2026-09-30