HLE gains appear to be slowing as GPT-5.6 and Claude 5 add only a few points
thegamebegins25 · reddit · 2026-07-25
A Reddit post argues that progress on Humanity’s Last Exam (HLE) may be slowing.
The author points out that:
- OpenAI’s GPT-5.6 blog post did not mention HLE;
- GPT-5.6 Sol (max) seems to beat GPT-5.5 (xhigh) by less than 3 points;
- Claude’s newest Fable/Opus/Sonnet 5 line also appears to improve by only about 3 points over Opus 4.8.
The post contrasts this with last year’s more frequent double-digit monthly gains and asks whether companies are now shifting focus to Agent’s Last Exam because HLE is getting harder to move.
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27