RSI Arena: 8 AI agents get 1,000 GPU-hours each to train a better model live
my_cat_can_code · x · 2026-09-29
RSI Arena opens live at COLM 2026: eight AI agents start from the same base model (Nemotron 3.5 Lightning 30B-A3B), each with $300 of API credit and 1,000 GPU-hours on 64× RTX PRO 6000 across 8 Slurm clusters, and get 144 hours to pick data, define benchmarks, and run experiments autonomously.
- Stage 1 kicks off Sept 29, 5 PM PT, streamed at rsiarena.live with a prize round for predicting the top three
- Stage 2 lets humans judge the models in the arena; agents keep training on that feedback with another 500 GPU-hours
- Checkpoints and data will be open-sourced, testing one concrete step of recursive self-improvement
More from coding & agent
- 50 researchers drop 303-page guide: RLVR lets small open models match giants on coding — mdancho84 · 2026-09-29
- If AI Coding Agents Get Reliable Enough, Frameworks May Disappear Entirely — bendee983 · 2026-09-29
- Langfuse open-sources MCP Server to debug agent runs without raw telemetry — blaizedsouza · 2026-09-29
- The Claude Code creator's 30-min vibe-coding workshop beats 100 YouTube guides — HeyAmit_ · 2026-09-29
- One Prompt, One Chat: Generating a Brand Motion-Graphics Video with Claude Code — TawohAwa · 2026-09-29
- The overlooked async RL curation pitfall: task runtime differences skew sampling — auto_grad_ · 2026-09-29