CAIS ships HLE-Rolling, a continuously updated fork of Humanity's Last Exam
devindkim · x · 2026-09-18
The CAIS team announced HLE-Rolling (cais/hle-rolling on Hugging Face), a dynamic fork of Humanity's Last Exam, updated as of September 17:
- Continuously cleaned using community feedback and external expert review, with easy questions swapped for harder ones from a held-out set
- Designed as a seamless migration path once frontier models hit the noise ceiling on the original HLE
- MIT-licensed, includes a canary string for training-data filtering; public redistribution is discouraged to protect benchmark integrity
- The team invites reports of bad questions
Teams benchmarking on HLE should consider switching to this cleaner version.
More from Research
- New academic spam: single-authored papers cold-pitching ARR service contributions — anmarasovic · 2026-09-18
- Index pretraining lifts humanoid zero-shot success from 8% to 56% — coreylynch · 2026-09-18
- A 'Life Diary' Eval Could Be the Toughest Test Yet for Continual Learning in LLMs — JohnnyNi13 · 2026-09-18
- Pretraining on Human Behavior Reportedly Boosts Task Success from 9% to 56% — Dr_Singularity · 2026-09-18
- Index pretraining lifts Helix 2.5 zero-shot success from 8% to 56%, generating 50 min of data per second — coreylynch · 2026-09-18
- LLMs got good at text and stayed bad at tables — and it's not just a training-data problem — FamiliarSlide7685 · 2026-09-18