All AI benchmarks are 'broken or saturated', new long-running loop eval launching next week
bindureddy · x · 2026-09-04
Bindu Reddy (Abacus.AI) says current AI benchmarks are 'totally broken or saturated' because they fundamentally don't test real-world long-running loops. She announces a new benchmark launching next week that fixes this, arguing evals need a revamp as models improve.
More from Models
- Matt Shumer reviews GPT-6 Astra: first model he trusts to run his inbox and business — mattshumer_ · 2026-09-04
- Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking — connoraxiotes · 2026-09-04
- Qwen 3.8 27B vs 3.6: quality up 8% but runtime 5x longer and 4x more tokens — DerTomsn · 2026-09-04
- Leak claims GPT-6 Astra trained on 100,000+ GPUs at OpenAI's Stargate site — BLUECOW009 · 2026-09-04
- Gary Marcus on GPT-6 Astra: symbolic world models vindicated, but not AGI — Gary Marcus · 2026-09-04
- Users dispute credit burn; provider says KV cache was always on, scaling across providers — arthurcolle · 2026-09-04