Will Depue says dangerous AI evals should be kept private, not published as numbers

willdepue · x · 2026-07-26

Will Depue argues that dangerous capability evals should usually stay private and be reported only as low/medium/high, not with public numbers.

He says the current “number go up” culture makes it too easy to optimize against published benchmarks. In the quoted reply, the point is framed more broadly as a warning that public evals can accelerate gain-of-function style research and make dangerous capabilities easier to measure and improve.

Related event: Musk and AI Safety Experts Call to Stop Public Dangerous Capability Evals(6 posts)→

Original post →

More from Safety

Safety channel →