Which benchmarks are still far from saturation? RLI tops out at 20%

MaximumIntention · reddit · 2026-09-28

A Reddit thread compiles benchmarks where frontier models still score very low (≤30%) and that are actively maintained: RLI (top score 20%) and ProgramBench (4.5%). Others like FormulaOne (0% on hardest set) and Esolang-bench (4.2%) appear no longer updated. As models improve rapidly, truly unsaturated benchmarks are becoming rare.

Original post →

More from Research

Research channel →