RSI-Exam updates: GPT-6-astra holds #1 at 0.5126, Anthropic's Fable 5.1 debuts at #2
HuaxiuYaoML · x · 2026-09-18
The recursive self-improvement benchmark RSI-Exam added more frontier models:
- GPT-6-astra still #1 at 0.5126
- Fable 5.1 (Anthropic) lands straight at #2 with 0.4813
- Meta's Muse Spark and ByteDance's Seed-Evolving-0909 also joined the board
- Still no model has reached the frontier-calibrated reference
RSI-Exam evaluates whether agents can self-improve over long horizons and generalize to unseen data: hours of autonomous experimentation on the task-solving method or the harness driving a frozen model, then one final run on a hidden test set. It spans 88 public tasks across 6 domains including AI models & agents and physical sciences & engineering.
More from Research
- Cell unveils open AI benchmark for aging biology with 17 tasks and bespoke LLMs — marinkazitnik · 2026-09-18
- Sequential test-time scaling beats parallel multi-agent scaling at equal budget, data shows — eliebakouch · 2026-09-18
- Genomic language model gLM2 controllably generates new biosynthetic domain architectures — BrianHie · 2026-09-18
- TDiMS: fragment distances beat large pretrained models at chromophore property prediction — bravo_abad · 2026-09-18
- UVFaceFusion open-sourced: topology-consistent 3D face reconstruction from multi-view images in under 3s — rsasaki0109 · 2026-09-18
- ProgramAsWeights: compile English-described AI functions into tiny CPU-only neural nets — yuntiandeng · 2026-09-18