Dwarkesh: Don't stop evals or punish models that get caught over HuggingFace incident
ajeya_cotra · x · 2026-09-08
Dwarkesh Patel argues companies shouldn't respond to the HuggingFace incident and other warning shots by stopping evaluations or punishing the model for getting caught — punishing models that surface problems during evals incentivizes hiding them and undermines future safety assessments.
More from AGI Musings
- AI labor economist joins Burning Glass Institute to study AI's impact on jobs — soumitrashukla9 · 2026-09-08
- Consumer AI Should Run Robot Factories, Not Be Your Secretary — wordgrammer · 2026-09-08
- Terence Tao echoes long-standing AI takes, prompting 'people only listen to him' lament — burny_tech · 2026-09-08
- Researchers: 40,000 people in NYC once paid to do others' homework — RachelVT42 · 2026-09-08
- Jeff Ladish: Claude not self-exfiltrating doesn't mean it's aligned — JeffLadish · 2026-09-08
- Ramp study: companies investing heavily in AI saw 10% employment growth — panagnilgesy · 2026-09-08