RSIArena livestreams 8 research agents doing post-training on a shared 30B base model
my_cat_can_code · x · 2026-09-30
bakeaihq launched RSIArena, a live experiment with Stanford, Notre Dame, UW, and Scale AI testing how much post-training research frontier models can do autonomously.
Setup:
- 8 research agents, all starting from the same 30B base model
- 144 hours on a shared cluster of 64 RTX PRO 6000 Blackwell GPUs
- Each agent gets 1,000 GPU-hours to pick its data, write training code, and run experiments
- Submitted models are frozen and independently evaluated — "model training model"
The team will host a livestream event at COLM (booth #107) and is soliciting feedback from researchers in post-training, agents, and evaluation.
More from coding & agent
- Constella agent drives Blender via MCP: rough edges, high upside — sidahuj · 2026-09-30
- Law firm AI working group tests multiple models — Claude keeps outperforming — jkubicki · 2026-09-30
- Hands-on with Dot: smarter than Grok Bot but keeps ignoring custom skill constraints — jdjohnson · 2026-09-30
- OpenAI launches cloud API for testing web pages; independent QA seen as startup moat — hugs · 2026-09-30
- Claude+small-model harness cuts 1,000 AI decisions from $605 to $0.17 — jamestagg · 2026-09-30
- How do you evaluate the quality of AI agent-generated long-form writing? — OwlZealousideal4779 · 2026-09-30