AI coding benchmark findings disputed over flawed data
A hands-on ranking based on Artificial Analysis data suggested open-source models need 10-20x tokens for interactive HITL coding, favoring closed models. Sergey Karayev disputed the finding, arguing the underlying benchmarks like deepswe-bench are narrow and impractical.
2026-09-01 ~ 2026-09-01 · 2 related posts
- Benchmark Report: OSS Models Need 10-20x Tokens, Slow for HITL Coding — nateberkopec · 2026-09-01
- Critique of AI Coding Benchmarks: Sparse Coverage and Low Utility — sergeykarayev · 2026-09-01