Benchmarks Deemed 'Cooked' as AI Models Rapidly Top Leaderboards

Anthropic's cheaper Haiku 5.5 outscored the pricier Opus 5 on multiple benchmarks, while Terminal-Bench was declared 'cooked,' fueling criticism that AI benchmarks are losing meaning as models saturate them faster than they can be updated.

2026-10-08 ~ 2026-10-08 · 2 related posts