antirez rips top coding benchmarks: mostly 'crap' testing tool friction, not real ability
antirez · x · 2026-09-03
Redis creator antirez argues that benchmarks often fail to align with real-world model behavior. He urges people to open a few tests from a top coding benchmark used by labs: most are, in his words, "crap" — stressing tool usage and friction between languages like C++ and Python more than genuine programming skill.
Related event: Redis Creator antirez Slams Coding Benchmarks as Mostly Garbage(2 posts)→
More from coding & agent
- Only 3 pure software businesses at the limit: agent-useful APIs, foundation models, autoresearch swarms — menhguin · 2026-09-03
- Cola Skill aggregates community-curated agent skills, with new picks daily — oran_ge · 2026-09-03
- Hacking llama.cpp to hot-swap knowledge into Qwen's Ngram PLE table — ortegaalfredo · 2026-09-03
- Qwen Releases E-Commerce Bench: Agents Run Stores for 365 Days, Few Learn — Alibaba_Qwen · 2026-09-03
- Three weeks of llama.cpp optimizations: collecting best t/s for Qwen3.8-27B — pmttyji · 2026-09-03
- Agents Don't Need to Talk: Watermarks, Git and Logs Are Their Hidden Channel — labeveryday · 2026-09-03