Benchmarks Deemed 'Cooked' as AI Models Rapidly Top Leaderboards
Anthropic's cheaper Haiku 5.5 outscored the pricier Opus 5 on multiple benchmarks, while Terminal-Bench was declared 'cooked,' fueling criticism that AI benchmarks are losing meaning as models saturate them faster than they can be updated.
2026-10-08 ~ 2026-10-08 · 2 related posts
- Haiku 5.5 Beats Opus 5 on GDPval While 75% Cheaper—'Meaningless Benchmarks,' Devs Joke — rickasaurus · 2026-10-08
- "Terminal-Bench is cooked": coders mock another saturated AI coding benchmark — rickasaurus · 2026-10-08