GPT-6.1 sol reportedly tops pure reasoning on HLE-Diamond at 67.4% using just 4k tokens
haider1 · x · 2026-09-30
X user haider1 shared (unofficial) HLE-Diamond results for GPT-6.1 sol:
- Overall: 6.1 sol at 53.2%, vs Fable 5.1 at 50.7% and Opus 5.5 at 54.6%
- On pure reasoning, 6.1 sol leads both: 67.4% vs 60.8% for Fable 5.1 and 62.4% for Opus 5.5
- Token efficiency is the standout: 4k tokens, compared to 16k for Fable and 10k for Opus
If accurate, the model matches or beats competitors while consuming a quarter to half their tokens. Third-party numbers, not yet officially confirmed.
More from Models
- ChatGPT Pro users report 6-Pro web chats capped at 100 per week — triestdain · 2026-09-30
- GPT-6.1 Sol fixes 44 of 105 planted bugs for $6.56, matching Astra at a fraction of the cost — PawelHuryn · 2026-09-30
- 5 months after Mythos Preview panic, an open model already matches it — mariofilhoml · 2026-09-30
- Artificial Analysis hosts first Seoul event on benchmarking and cost-per-task AI — ArtificialAnlys · 2026-09-30
- GPT 6.1 Sol Only +3 on BridgeBench, 82 Points Behind Astra: Benchmaxing Suspected — RexDouglass · 2026-09-30
- Anthropic's Latest Blog Post Mentions Zhipu's GLM 5.3 — gnukeith · 2026-09-30