RSI-Exam lands in State of AI report: top model Opus 5.5 scores just ~0.53

cihangxie · x · 2026-10-09

The RSI-Exam benchmark by Huaxiu Yao's team has been included in this year's State of AI report. The team argues that as AI R&D benchmarks saturate, evals with real headroom are needed: RSI-Exam contains 88 RSI tasks, and the top model (Opus 5.5) still scores only 0.53, highlighting how far frontier models remain from autonomous AI research capability.

Original post →

More from Research

Research channel →