antirez: coding benchmarks are 'mostly crap' and badly misaligned with how models train
antirez · x · 2026-09-03
Redis creator antirez argues that training progress has little to do with the benchmarks later used to evaluate models, and bets every major AI lab keeps a private benchmark set to truly gauge model strength. He urges people to open the top coding benchmarks labs rely on: they are 'mostly crap', stressing tool use and friction between C++, Python and the like rather than real-world model behavior.
Related event: Redis Creator antirez Slams Coding Benchmarks as Mostly Garbage(2 posts)→
More from Models
- Ethan Mollick tests Gemini 3.8 Flash: fast but no match for Fable 5.1 on shaders — eldonredwards · 2026-09-03
- Meta's Spark 1.3 nears frontier and its scorched-earth pricing may threaten AI labs, argues investor — RihardJarc · 2026-09-03
- Claude Fable 5.1 (high) hits 92.3% on WeirdML, beating Fable 5 by 0.4% for new SOTA — teortaxesTex · 2026-09-03
- Would OpenAI bet $100M+ on Looped Transformer for Astra without scaling proof? — teortaxesTex · 2026-09-03
- Enterprises pay 10x-20x more to keep data out of AI training, suggesting a routing-layer startup — random_walker · 2026-09-03
- Hermes and Alibaba Cloud demo Qwen 3.8 for 200+ AI founders at largest joint event — Teknium · 2026-09-03