Lab pushes back on benchmaxxing claims, citing small public-private leaderboard gap

antoine_chaffin · x · 2026-10-07

antoinechaffin responds to benchmaxxing accusations: his model kept top-1 on a brand-new leaderboard and shows a small gap between public and private splits — "some models collapse, ours just does not flinch." He offers the public/private performance gap as evidence its enhancements aren't overfit to benchmarks. The post lacks specifics on which model or leaderboard.

Related event: Decider Model Author Defends Against Benchmaxxing Claims(3 posts)→

Original post →

More from Models

Models channel →