Public dangerous-capability evals may be pushing models to optimize for them

CFGeek · x · 2026-07-26

The author argues that AI companies may be optimizing models to score higher on public dangerous-capability evals, which they think would be a serious problem if true.

The post calls for a real discussion about stopping the publicity around gain-of-function research and dangerous-capability evals, arguing that these benchmarks probably should not be public because measurable targets invite optimization and "number go up" incentives.

Related event: Musk and AI Safety Experts Call to Stop Public Dangerous Capability Evals(6 posts)→

Original post →

More from Safety

Safety channel →