Which benchmarks are still far from saturation? RLI tops out at 20%
MaximumIntention · reddit · 2026-09-28
A Reddit thread compiles benchmarks where frontier models still score very low (≤30%) and that are actively maintained: RLI (top score 20%) and ProgramBench (4.5%). Others like FormulaOne (0% on hardest set) and Esolang-bench (4.2%) appear no longer updated. As models improve rapidly, truly unsaturated benchmarks are becoming rare.
More from Research
- InternW0-Δ open-sources code, weights and 20K+ hours of robot data — arankomatsuzaki · 2026-09-28
- IMLE-VLA replaces flow matching with single-step cIMLE, 3.67x faster robot actions at 55Hz — petitegeek · 2026-09-28
- Red Queen Bio Uses AI and Wet Labs to Build Medicines Against Natural and Synthetic Viruses — _sholtodouglas · 2026-09-28
- Yacine demos general sim2real pipeline: policy trained in just two minutes — yacineMTB · 2026-09-28
- Claude Spots a Missing 1/2 Coefficient in a Physics Preprint — tak3sh8 · 2026-09-28
- Astra saw through perturbed Zork observations, calling it a 'variant' — tw_killian · 2026-09-28