AI safety researcher warns that public dangerous-capability evals can become hill-climbable

willdepue · x · 2026-07-26

The post argues that frontier labs and researchers should measure dangerous capabilities, but that the current approach is getting this wrong.

Key points:

The thread also cites the recent Hugging Face hack as a reminder that harness/eval engineering is non-trivial and that dangerous capability research is not being handled carefully enough.

Related event: Musk and AI Safety Experts Call to Stop Public Dangerous Capability Evals(6 posts)→

Original post →

More from Safety

Safety channel →