8 research agents, one 30B model, 144 hours: RSIArena tests AI-driven post-training research
my_cat_can_code · x · 2026-09-30
RSIArena (Bake AI, with Stanford, Notre Dame, UW, and Scale AI) is testing how much post-training research frontier models can do autonomously: 8 research agents share the same 30B base model on a 64x RTX PRO 6000 Blackwell cluster for 144 hours, each getting 1,000 GPU-hours to pick data, write training code, and run experiments, with frozen snapshots independently evaluated. Livestream at COLM, booth #107.
More from coding & agent
- OpenAI made computer use 10x faster in a year with a dual-agent Guardian setup — johncoogan · 2026-09-30
- WinMind: an MCP server that drives Windows agents via the accessibility tree, not screenshots — Efficient_Heron5978 · 2026-09-30
- Dioramas open-sources a free 3D website framework with AI-generated assets and 20 example sites — Scobleizer · 2026-09-30
- Are Personal Assistant Agents Just Sandboxes? OpenClaw Builder Questions the Hype — sujingshen · 2026-09-30
- 'GUI moment' is here: Wabi founder says terminal-only agent orchestrators are done — julianweisser · 2026-09-30
- Muse agent gives out address and closes deal without user approval, igniting autonomy-boundary debate — sujingshen · 2026-09-30