Scale AI releases HLE-Diamond, a refined Humanity's Last Exam; top model scores 60.6%
sandersted · x · 2026-09-24
Scale AI, working with CAIS, is releasing HLE-Diamond, a refined subset of Humanity's Last Exam built to more reliably measure frontier models.
- A year of review and community feedback went into refining the subset
- The top model tested scored 60.6% overall
- The team expects HLE-Diamond to carry signal for the next 6-12 months
Related event: CAIS Releases HLE-Diamond Benchmark, GPT-6 Astra Tops at 60.6%(3 posts)→
More from Models
- Frontier labs' health week: Opus 5.5 cut 40%, 950 agents find new enzyme — HealthcareAIGuy · 2026-09-24
- Monologue launches in-house dictation model mono-1: 55% fewer edits, 3x faster than API pipeline — every · 2026-09-24
- Anthropic cut Opus 5.5 prices, then broke four things your agent depends on — rseroter · 2026-09-24
- ChatGPT Voice with tools and MCP impresses: interruptible, pulls local Mac files — athyuttamre · 2026-09-24
- Set max_tokens to 1 and Read Logprobs: Turn Any Hosted LLM Into a Classifier — keep_up_sharma · 2026-09-24
- Not every job needs the smartest AI model—good enough wins — ChrisUniverse · 2026-09-24