RSI-Exam lands in State of AI report: top model Opus 5.5 scores just ~0.53
cihangxie · x · 2026-10-09
The RSI-Exam benchmark by Huaxiu Yao's team has been included in this year's State of AI report. The team argues that as AI R&D benchmarks saturate, evals with real headroom are needed: RSI-Exam contains 88 RSI tasks, and the top model (Opus 5.5) still scores only 0.53, highlighting how far frontier models remain from autonomous AI research capability.
More from Research
- Open differential geometry book by Pinkall & Gross joins ChapterPal with an AI tutor — burkov · 2026-10-09
- KAIST's ME-World tackles multi-agent egocentric world modeling with joint denoising — kaist-ai · 2026-10-09
- Controlled study: LLMs struggle to recover latent sequential structure despite long context — illinois · 2026-10-09
- TerraVis quantifies world-grounded visual consistency failures in text-to-image models — the-aiml · 2026-10-09
- OneSearch-VL: unified multimodal deep research agent beats Qwen3-VL by 20 points — Hongyu Li · 2026-10-09
- SpaceCast-Bench: best VLM hits 58.0% on predictive spatial reasoning vs 87.2% human — zju · 2026-10-09