HLE gains appear to be slowing as GPT-5.6 and Claude 5 add only a few points
thegamebegins25 · reddit · 2026-07-25
A Reddit post argues that progress on Humanity’s Last Exam (HLE) may be slowing.
The author points out that:
- OpenAI’s GPT-5.6 blog post did not mention HLE;
- GPT-5.6 Sol (max) seems to beat GPT-5.5 (xhigh) by less than 3 points;
- Claude’s newest Fable/Opus/Sonnet 5 line also appears to improve by only about 3 points over Opus 4.8.
The post contrasts this with last year’s more frequent double-digit monthly gains and asks whether companies are now shifting focus to Agent’s Last Exam because HLE is getting harder to move.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11