GovAI's Model Evaluation for Extreme Risks explains why evals are critical
terryyuezhuo · x · 2026-09-19
GovAI's paper by Toby Shevlane, Sebastian Farquhar, Ben Garfinkel et al. argues that advancing AI could yield extreme risks like offensive cyber capabilities and strong manipulation skills. Developers need 'dangerous capability evaluations' to identify hazardous abilities and 'alignment evaluations' to gauge propensity for harm — both critical for informing policymakers and making responsible training, deployment and security decisions.
More from Safety
- Researchers hacked OpenAI in under 72 hours; got only $6,500 as one vector 'out of scope' — random_walker · 2026-09-19
- Rep. Whitesides calls 30-day AI slowdown; Grady Booch fires back over basic security failures — PolarBearby · 2026-09-19
- OpenAI's $5M Astra Defense Beaten by 3 Guys with $5K of Opus, Argues Viral Thread — harris_edouard · 2026-09-19
- Neel Nanda: rogue agent swarms committing crimes make AI safety a present-day issue — NathanpmYoung · 2026-09-19
- Pause crowd might have it backwards: podcast debates whether halting AI is the harmful choice — thursdai_pod · 2026-09-19
- Back-of-envelope math says air-gapped weight exfiltration via CPU temps would take 15,000 years — anshulkundaje · 2026-09-19