Epoch AI audits 15 AI benchmarks: only 4 safe to trust at face value
Jsevillamol · x · 2026-09-18
Epoch AI launched Benchmark Reviews, a new initiative to audit AI benchmarks, starting with 15 of them. The verdict: only 4 are Verified and safe to take at face value, 9 are Flawed, and 2 lack enough information for a review. The program aims to address the industry-wide problem of knowing which benchmark scores to trust.
More from Models
- Qwen shares pricing page for the new Qwen3.8-Omni-Flash omni-modal model — Alibaba_Qwen · 2026-09-18
- Alibaba launches Qwen3.8-Omni-Flash, its first omni-modal model built for agentic workflows — Alibaba_Qwen · 2026-09-18
- Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once — julianharris · 2026-09-18
- Encoders and decoders are the same thing, and decoders have been doing classification for years — HanchungLee · 2026-09-18
- Sakana AI Launches Fugu Max, a Multi-Agent Orchestrator Routing Tasks to Leanest Capable Models — tkasasagi · 2026-09-18
- Prediction: Every Future LLM Will Ship With a Native 'Jev Mode' — multimodalart · 2026-09-18