RSI-Exam opens 35 of 88 tasks, hailed as the next SWE-bench for self-improving AI
cihangxie · x · 2026-09-12
HuaxiuYaoML's team announced that 35 of the 88 tasks in RSI-Exam are now public, with the benchmark continuously expanding and welcoming contributors from different research fields to submit challenging tasks. YC researcher YGandelsman called it one of the most exciting benchmarks to follow, "feels like the next SWE-bench." RSI-Exam targets models' recursive self-improvement abilities via frontier research problems.
More from Research
- Open-weights SUPlime beats pyannote's commercial diarization with 15.86 DER — solyarisoftware · 2026-09-12
- AI out-persuades world champion debaters, raises donations nearly 3x better than pros — ben_j_todd · 2026-09-12
- Shifting local token interactions from inference to training-time lookups seen as a scaling win — AccBalanced · 2026-09-12
- Open-source RL for large MoEs with zero train-infer mismatch, teaching Qwen3.6-35B-A3B to play Wordle — kastnerkyle · 2026-09-12
- ICLR 2026 paper MoM: multiple memory states fix linear models' recall weakness — kastnerkyle · 2026-09-12
- Harvard researcher turned a fruit fly brain connectome into a Bitcoin trading bot — Scobleizer · 2026-09-12