RSI-Exam Benchmark: Quantifying Recursive Self-Improvement in LLMs

cihangxie · x · 2026-08-27

RSI-Exam is a new benchmark designed to quantify Recursive Self-Improvement (RSI) capabilities in LLMs. Instead of measuring weight modifications, it tests if an agent can improve an 'executable work product' through hours of autonomous experimentation and verify generalization on a hidden test set. The benchmark includes 88 tasks (35 public, 53 private) spanning 6 domains like physical sciences, AI models, and systems. The leaderboard shows Claude and GPT forming a top tier, though there remains significant room for improvement.

Related event: RSI-Exam Benchmark Launches with 88 Tasks to Test Recursive Self-Improvement in AI(5 posts)→

Original post →

More from coding & agent

coding & agent channel →