RSI-Exam leaderboard update: GPT-6-astra tops recursive self-improvement benchmark at 0.5126
yuyinzhou_cs · x · 2026-09-19
- New leaderboard update for RSI-Exam, a benchmark testing whether agents can autonomously improve an existing method over long-horizon experiments and generalize to hidden data. Six models added.
- GPT-6-astra (OpenAI) takes #1 at 0.5126, Fable 5.1 (Anthropic) #2 at 0.4813, followed by Qwen3.8 Max (0.3923), Muse Spark 1.3 (0.3882), Gemini 3.8 Flash (0.3406), and Seed Evolving-0909 (0.3243).
- The benchmark spans 88 executable research tasks across 6 domains.
- No model has yet reached the frontier-calibrated reference score of 0.6 — huge headroom remains for recursive self-improvement.
Related event: GPT-6-astra Tops Updated RSI-Exam Leaderboard as Fable 5.1 Debuts Second(3 posts)→
More from Models
- Anthropic's Fable 5.1 lands in Kiro, built for long-running context-heavy agentic sessions — DigitalColmer · 2026-09-19
- Dev benchmark: Astra beats sol and fable on LOC, API cost and code quality in 10-15 feature test — robleclerc · 2026-09-19
- Free Jev on Vercel Is So Rate-Limited It's Barely Usable — tristanbob · 2026-09-19
- Jev hype signals strong appetite for more tokens/sec, but limits vs general LLMs — rbhar90 · 2026-09-19
- One compact operation in Claude desktop ate 20% of a 5-hour usage limit — _xjdr · 2026-09-19
- Producer says Suno now flags every self-made DAW track, blocking uploads — tetsuoai · 2026-09-19