Mollick: public AI benchmarks are broken and underestimate AI abilities
emollick · x · 2026-09-16
Ethan Mollick argues the state of public AI benchmarking is "dire" and undermining our ability to gauge current AI. Per the cited paper, most famous benchmarks are saturated, while the non-saturated ones are riddled with errors that vastly underestimate AI abilities.
Related event: Mollick: public AI benchmarks are broken and underestimate models(2 posts)→
More from Models
- Cohere launches Confidential Computing in Model Vault: encrypted even during inference — cohere · 2026-09-17
- llama.cpp merges support for DFM Mimir 1B, GGUF weights on Hugging Face — noctrex · 2026-09-16
- Researcher warns LLM evals are broken after reading 'truly horrible' model traces — IanArawjo · 2026-09-16
- OpenAI's 10,000-word essay on a self-evolving 'alien mind' sparks ownership clash with startup JoyIn — SarahAnnabels · 2026-09-16
- Jev is a smart classifier, not a frontier LLM — a contrarian analysis — DavideCrapis · 2026-09-16
- Custom Midtrain Plus Own RL Matches Astra Max at Half the Inference Price — hsu_byron · 2026-09-16