Qwen3.8 Max jumps 22% on new RSI-Exam benchmark for recursive self-improvement
cihangxie · x · 2026-09-05
RSI-Exam 0.1 is a new benchmark of 88 executable research tasks across 6 domains (virtual cells, TPU kernels, chip design, quant finance, agent harnesses, model distillation) testing whether frontier agents can turn a weak method into one that performs better on hidden data. Protocol: the agent iterates a method on visible data; only the final artifact is evaluated in a fresh container on the hidden set. Opus 5 currently leads the 88-task leaderboard, while Qwen3.8 Max-0902 jumped from 0.322 to 0.392 on recursive self-improvement tasks (22% improvement). Task contributions are open.
Related event: Qwen3.8-Max Gains 22% on RSI-Exam After Coding-Focused Training(2 posts)→
More from Models
- AI now writes 51% of Vercel design engineering code in August 2026, up from 50% — evilrabbit_ · 2026-09-05
- Meta models hit 43% share on opencode, up from 3.5% two weeks ago — zephyr_z9 · 2026-09-05
- Alexandr Wang Announces Muse Spark 1.3 with Max Reasoning and One-Line CLI Install — alexandr_wang · 2026-09-05
- Anthropic posts a complete Lean 4 machine-checked proof of Fermat's Last Theorem — scaling01 · 2026-09-05
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05