Will Depue says dangerous AI evals should be kept private, not published as numbers
willdepue · x · 2026-07-26
Will Depue argues that dangerous capability evals should usually stay private and be reported only as low/medium/high, not with public numbers.
He says the current “number go up” culture makes it too easy to optimize against published benchmarks. In the quoted reply, the point is framed more broadly as a warning that public evals can accelerate gain-of-function style research and make dangerous capabilities easier to measure and improve.
Related event: Musk and AI Safety Experts Call to Stop Public Dangerous Capability Evals(6 posts)→
More from Safety
- DHH Slams 'GDPR Is Good' Take: Vague Rules Birthed a Bureaucratic Beast — dhh · 2026-09-11
- Houthis tried to use Claude to design missile software, Anthropic says it blocked the attempts — Affectionate_Bee6434 · 2026-09-11
- AI safety community mocked as 'bridge engineers' who say bridges can never be safe — Dan_Jeffries1 · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11