CAIS releases HLE-Diamond, a refined 1,000-question HLE subset; GPT-6 Astra tops it at 60.6%

steipete · x · 2026-09-24

CAIS and Scale AI have released HLE-Diamond, a refined subset of Humanity's Last Exam distilled over a year of cleaning with input from research communities. It contains 1,000 questions — 500 reasoning and 500 expert-knowledge — designed to be answerable closed-book.

No-tools leaderboard (reasoning set to high):

With tools (web+code) scores jump sharply: GPT-6 Astra reaches 82.9%, Claude Opus 5.5 hits 73.9%, and Claude Opus 5 climbs to 69.1%. The team also published recommended settings for evaluating agentic systems on the benchmark.

Related event: CAIS Releases HLE-Diamond Benchmark, GPT-6 Astra Tops at 60.6%(3 posts)→

Original post →

More from Models

Models channel →