RSI Arena: 8 research agents autonomously train rival models for human judging
my_cat_can_code · x · 2026-09-30
LIVE RSI Arena, run with Stanford, Notre Dame, UW and Scale AI, tests how much post-training research frontier models can do autonomously: 8 research agents start from the same 30B base model, sharing 64 RTX PRO 6000 Blackwell GPUs for 144 hours (1,000 GPU-hours each), choosing their own data and writing training code. A final human evaluation decides whose trained model wins.
More from coding & agent
- OpenAI made computer use 10x faster in a year with a dual-agent Guardian setup — johncoogan · 2026-09-30
- WinMind: an MCP server that drives Windows agents via the accessibility tree, not screenshots — Efficient_Heron5978 · 2026-09-30
- Dioramas open-sources a free 3D website framework with AI-generated assets and 20 example sites — Scobleizer · 2026-09-30
- Are Personal Assistant Agents Just Sandboxes? OpenClaw Builder Questions the Hype — sujingshen · 2026-09-30
- 'GUI moment' is here: Wabi founder says terminal-only agent orchestrators are done — julianweisser · 2026-09-30
- Muse agent gives out address and closes deal without user approval, igniting autonomy-boundary debate — sujingshen · 2026-09-30