RSIArena pits 8 research agents on one 30B model with 64 RTX 6000 GPUs for 144 hours

my_cat_can_code · x · 2026-09-30

RSIArena, run with Stanford, Notre Dame, UW, and Scale AI, tests how much post-training research frontier models can do autonomously: 8 research agents share a 30B base model and a cluster of 64 RTX PRO 6000 Blackwell GPUs for 144 hours, each getting 1,000 GPU-hours to pick data, write training code, and run experiments. Submitted models are frozen and independently evaluated, with a livestream and a COLM booth for feedback.

Related event: RSIArena: 8 Research Agents Self-Study Post-Training for 144 Hours(2 posts)→

Original post →

More from coding & agent

coding & agent channel →