RSIArena Live Test: 8 Agents on the Same 30B Base Compete at Autonomous Post-Training Research

RSIArena (by Bake AI), together with Stanford, Notre Dame, UW and Scale AI, has launched a livestreamed experiment testing how much post-training research frontier models can complete on their own. Eight research agents, all starting from the same base model Nemotron 3.5 Lightning 30B-A3B, each carry out roughly 144 hours of recursive self-improvement (RSI) research on a shared pool of 64 RTX PRO 6000 Blackwell GPUs, with human judges deciding the winner. The experiment kicks off live at COLM 2026.

Confirmed

Why it matters

2026-09-29 ~ 2026-09-30 · 8 related posts

Primary sources

3 near-duplicate retellings: my_cat_can_code · my_cat_can_code · my_cat_can_code