Qwen3.8 Max jumps 22% on new RSI-Exam benchmark for recursive self-improvement

cihangxie · x · 2026-09-05

RSI-Exam 0.1 is a new benchmark of 88 executable research tasks across 6 domains (virtual cells, TPU kernels, chip design, quant finance, agent harnesses, model distillation) testing whether frontier agents can turn a weak method into one that performs better on hidden data. Protocol: the agent iterates a method on visible data; only the final artifact is evaluated in a fresh container on the hidden set. Opus 5 currently leads the 88-task leaderboard, while Qwen3.8 Max-0902 jumped from 0.322 to 0.392 on recursive self-improvement tasks (22% improvement). Task contributions are open.

Related event: Qwen3.8-Max Gains 22% on RSI-Exam After Coding-Focused Training(2 posts)→

Original post →

More from Models

Models channel →