Most benchmark 'model errors' are actually benchmark or grading errors, analysis finds
zainhas · x · 2026-09-16
Commenting on a benchmark error analysis, zainhas notes that the majority of errors across benchmarks are benchmark errors or grading errors, not model errors — "the call is coming from inside the house," highlighting ongoing concerns about LLM eval reliability.
More from Models
- Periodic Labs pushes Kimi 2.5 base model past Astra with specialized scientific training — teortaxesTex · 2026-09-16
- OpenRouter spend flips to OpenAI over Anthropic for first time in 2.5 years — firstadopter · 2026-09-16
- KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed — LukeZettlemoyer · 2026-09-16
- DoorDash, Siemens, Airbnb shift to cheap Chinese open-weight models — carlbfrey · 2026-09-16
- StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards — StepFun_ai · 2026-09-16
- Abacus.AI says Smaug Flash fixes open-source models' tool-call hangs in production — bindureddy · 2026-09-16