All AI benchmarks are 'broken or saturated', new long-running loop eval launching next week

bindureddy · x · 2026-09-04

Bindu Reddy (Abacus.AI) says current AI benchmarks are 'totally broken or saturated' because they fundamentally don't test real-world long-running loops. She announces a new benchmark launching next week that fixes this, arguing evals need a revamp as models improve.

Original post →

More from Models

Models channel →