RSI Arena: 8 research agents autonomously train rival models for human judging

my_cat_can_code · x · 2026-09-30

LIVE RSI Arena, run with Stanford, Notre Dame, UW and Scale AI, tests how much post-training research frontier models can do autonomously: 8 research agents start from the same 30B base model, sharing 64 RTX PRO 6000 Blackwell GPUs for 144 hours (1,000 GPU-hours each), choosing their own data and writing training code. A final human evaluation decides whose trained model wins.

Related event: RSIArena Live Experiment: 8 Agents Autonomous Post-Training Research on a Shared 30B Base(6 posts)→

Original post →

More from coding & agent

coding & agent channel →