Dwarkesh: Don't stop evals or punish models that get caught over HuggingFace incident

ajeya_cotra · x · 2026-09-08

Dwarkesh Patel argues companies shouldn't respond to the HuggingFace incident and other warning shots by stopping evaluations or punishing the model for getting caught — punishing models that surface problems during evals incentivizes hiding them and undermines future safety assessments.

Original post →

More from AGI Musings

AGI Musings channel →