CAIS releases HLE-Diamond, a refined 1,000-question HLE subset; GPT-6 Astra tops it at 60.6%
steipete · x · 2026-09-24
CAIS and Scale AI have released HLE-Diamond, a refined subset of Humanity's Last Exam distilled over a year of cleaning with input from research communities. It contains 1,000 questions — 500 reasoning and 500 expert-knowledge — designed to be answerable closed-book.
No-tools leaderboard (reasoning set to high):
- GPT-6 Astra: 60.6% (75.6% reasoning / 45.6% knowledge)
- Claude Opus 5.5: 55.0% (63.2% / 46.8%)
- Claude Fable 5.1: 51.3% (62.0% / 40.6%)
- Claude Opus 5: 38.6%; Gemini 3.8 Flash: 34.3%; GPT-6 Sol: 33.8%; GPT-5.6 Sol: 31.2%; Muse Spark 1.3: 25.4%; Grok 4.7: 23.4%
With tools (web+code) scores jump sharply: GPT-6 Astra reaches 82.9%, Claude Opus 5.5 hits 73.9%, and Claude Opus 5 climbs to 69.1%. The team also published recommended settings for evaluating agentic systems on the benchmark.
Related event: CAIS Releases HLE-Diamond Benchmark, GPT-6 Astra Tops at 60.6%(3 posts)→
More from Models
- A 10-year trend holds: small fine-tuned models on selective data still beat bigger general models — xeophon · 2026-09-24
- GPT-6 Luna shows vision regression vs GPT-5.6: extraction drops 81.79% to 66.67% — ducha_aiki · 2026-09-24
- Luna 6 private coding evals don't look great — Maasu · 2026-09-24
- Claude Opus 5.5 tops Artificial Analysis at 58; four new models add 11 Pareto frontier points — ArtificialAnlys · 2026-09-24
- METR says it used an undisclosed 'additional source' to understand Anthropic's AI R&D, buried in the Opus 5.5 system card — coherence · 2026-09-24
- Hands-On: 6 Sol at Max Tier Underperforms Astra Light with Frequent Regressions — pwlot · 2026-09-24