antirez: coding benchmarks are 'mostly crap' and badly misaligned with how models train

antirez · x · 2026-09-03

Redis creator antirez argues that training progress has little to do with the benchmarks later used to evaluate models, and bets every major AI lab keeps a private benchmark set to truly gauge model strength. He urges people to open the top coding benchmarks labs rely on: they are 'mostly crap', stressing tool use and friction between C++, Python and the like rather than real-world model behavior.

Related event: Redis Creator antirez Slams Coding Benchmarks as Mostly Garbage(2 posts)→

Original post →

More from Models

Models channel →