RSI-Exam updates: GPT-6-astra holds #1 at 0.5126, Anthropic's Fable 5.1 debuts at #2

HuaxiuYaoML · x · 2026-09-18

The recursive self-improvement benchmark RSI-Exam added more frontier models:

RSI-Exam evaluates whether agents can self-improve over long horizons and generalize to unseen data: hours of autonomous experimentation on the task-solving method or the harness driving a frozen model, then one final run on a hidden test set. It spans 88 public tasks across 6 domains including AI models & agents and physical sciences & engineering.

Original post →

More from Research

Research channel →