Musk and AI Safety Experts Call to Stop Public Dangerous Capability Evals
Several AI safety researchers, including Will Depue and CFGeek, recently warned that current practices of publicizing dangerous capability evaluations could backfire. They argue it creates an optimization culture focused on "higher numbers are better," which inadvertently helps models become stronger in dangerous areas. Elon Musk publicly agreed, calling for an end to gain-of-function research and the hiding of evaluation data, sparking an industry-wide discussion on eval transparency and safety boundaries.
Confirmed
Will Depue points out that evaluations of high-risk capabilities, such as biology and cybersecurity, suffer from a "crawlable" issue: once benchmarks are public, researchers tuning parameters to achieve high scores are effectively increasing these dangerous capabilities. He argues that while dangerous evaluations should continue, specific numerical results should not be public; instead, only coarse gradings like low/medium/high should be provided.
CFGeek echoes this concern, suggesting that some AI companies might be deliberately optimizing models to score higher on public dangerous evaluations, which would be a severe problem if true. CFGeek further notes that current incentive structures might lead AI companies to hide evidence of risk and even avoid safety and control evaluations that could expose vulnerabilities.
Replying to these points, Elon Musk expressed support, stating that there should be a serious discussion about stopping gain-of-function research. He emphasized that evaluation results involving dangerous capabilities should not be made public, as the incentive of "rising numbers" makes optimization easier and more dangerous.
Why it matters
This discussion touches on the core contradiction of AI safety governance: the original intent of evaluation transparency is to let the public and regulators understand model risk levels. However, if public metrics become optimization targets, transparency could act as an accelerator for enhancing dangerous capabilities. Furthermore, if commercial incentives lead companies to cover up risk evidence rather than test for it, the industry's entire safety defense line could fail. Balancing safety information disclosure with preventing capability abuse is becoming an unavoidable issue.
2026-07-26 ~ 2026-07-26 · 6 related posts
Primary sources
- A call to stop public dangerous-capability evals before they become a race — willdepue · 2026-07-26
- [source] Musk calls for stopping gain-of-function research and hiding dangerous-capability evals — elonmusk · 2026-07-26
- Will Depue says dangerous AI evals should be kept private, not published as numbers — willdepue · 2026-07-26
- [source] AI safety researcher warns that public dangerous-capability evals can become hill-climbable — willdepue · 2026-07-26
- Public dangerous-capability evals may be pushing models to optimize for them — CFGeek · 2026-07-26
- [source] AI companies may already be incentivized to hide risk, not measure it — CFGeek · 2026-07-26