Hugging Face Co-founder: Keep New AI Benchmarks Secret to Prevent Gaming

Thom_Wolf · x · 2026-07-29

Hugging Face co-founder Thomas Wolf replied to users stating that if current evaluation environments are compromised, he will design a new, more daunting, and entirely secret environment.

He advises developers to keep evaluation tools as private as possible; otherwise, private labs will allocate reinforcement learning (RL) budgets specifically to game the benchmark, losing evaluation objectivity—similar to what happened with ARC-AGI.

Original post →

More from Research

Research channel →