Claude Opus 5.5 Tops RSI-Exam at 0.536, Edging Out GPT-6 on 88 Research Tasks
HuaxiuYaoML · x · 2026-10-02
- RSI-Exam, a new benchmark for recursive self-improvement, evaluates whether AI agents can improve themselves over long autonomous runs (hours of experimentation) and generalize to a hidden test set across 88 executable research tasks in 6 domains.
- Claude Opus 5.5 (Anthropic) takes #1 with 0.5357, leading on full, public, and private splits.
- GPT-6-astra is second at 0.5126.
- Grok 4.7 debuts at 0.3842, up from Grok 4.6's 0.3671.
- Gemini 4 is not yet on the board, with the poster teasing it's "on the way".
More from Models
- Adaptive effort is likely the killer app for dynamic looped transformers — willcb · 2026-10-02
- Dev says Claude and Claude Code are getting worse day by day — Pavan_Belagatti · 2026-10-02
- Abliterated Large V2 lands on Venice: refusal-free AI for red teamers, anonymously — 0xAllen_ · 2026-10-02
- GLM 5.3 lands in Cursor, with GLM 5.3 Max topping CursorBench 4.0 among open-weight models — zainhas · 2026-10-02
- Banned Anthropic users can't get refunds on $200 plans, dev calls for a refund guide — lxfater · 2026-10-02
- Both Columns, One Vendor: What Opus 5.5 Benchmarks Can and Cannot Show — maier_ak · 2026-10-02