RSIArena Live Test: 8 Agents on the Same 30B Base Compete at Autonomous Post-Training Research
RSIArena (by Bake AI), together with Stanford, Notre Dame, UW and Scale AI, has launched a livestreamed experiment testing how much post-training research frontier models can complete on their own. Eight research agents, all starting from the same base model Nemotron 3.5 Lightning 30B-A3B, each carry out roughly 144 hours of recursive self-improvement (RSI) research on a shared pool of 64 RTX PRO 6000 Blackwell GPUs, with human judges deciding the winner. The experiment kicks off live at COLM 2026.
Confirmed
- The experiment is named RSIArena, led by bakeaihq (Bake AI) with partners including Stanford, Notre Dame, UW and Scale AI
- 8 research agents, all built on the same 30B base model (Nemotron 3.5 Lightning 30B-A3B)
- Compute setup: 64 shared RTX PRO 6000 Blackwell GPUs, with 1000 GPU hours per agent
- Each agent also gets a $300 API budget
- The experiment runs as a livestream for about 144 hours, as a real-world test of recursive self-improvement (RSI)
- Winners are determined by human judges
- The experiment kicks off on-site at COLM 2026
Why it matters
- It is a large-scale, publicly livestreamed test of the core question of whether frontier models can independently carry out post-training research, and having 8 agents compete on the same task with the same resources enables direct comparison of autonomous research capabilities
- Resources per participant (1000 GPU hours, $300 API budget) are explicitly capped, making results comparable under controlled costs and providing a reproducible reference framework for evaluating recursive self-improvement
2026-09-29 ~ 2026-09-30 · 8 related posts
Primary sources
- [source] RSI Arena: 8 AI agents get 1,000 GPU-hours each to train a better model live — my_cat_can_code · 2026-09-29
- 8 research agents self-train a 30B model for 144 hours in RSIArena livestream experiment — my_cat_can_code · 2026-09-30
- [source] 8 research agents, one 30B model, 144 hours: RSIArena tests AI-driven post-training research — my_cat_can_code · 2026-09-30
- RSI Arena: 8 research agents autonomously train rival models for human judging — my_cat_can_code · 2026-09-30
- Stanford, Scale AI and others to livestream research agents doing autonomous post-training at COLM 2026 — my_cat_can_code · 2026-09-30
3 near-duplicate retellings: my_cat_can_code · my_cat_can_code · my_cat_can_code