OpenRSI-Index v0.1 launches: a 1k-GPU benchmark for recursive self-improvement with 60-hour agent runs
RulinShao · x · 2026-09-24
OpenRSI, co-led by MIT-IBM Watson AI Lab and Amazon A-EVO Lab with advisors including Yejin Choi, Jianfeng Gao, and Wenhu Chen, released OpenRSI-Index v0.1, an open benchmark measuring whether AI research agents can recursively improve foundation-model development beyond human baselines:
- Turns open-source projects into autoresearch environments on production-scale clusters (1k GPUs), with agent trajectories lasting 60+ hours; building v0.1 took 100K+ H100-hours
- Signature tasks include Marin-Scaling-Ladder pretraining (18,432 H100-hours per run, 550M→2.545B), Qwen-122B-RL-Merge post-training (20,864 H100-hours), and GPIC Leaderboard vision-gen (5,815 H100-hours)
- Partners include University of Washington, UC Berkeley, Texas A&M, and NUS; the team invites task and compute contributors
Related event: OpenRSI Launches Open Benchmark for Recursive Self-Improvement(2 posts)→
More from Research
- New DAYJOB benchmark: best agents complete only ~25% of real knowledge work — echen · 2026-09-24
- OpenAI Releases MentalHealthBench, an Open Benchmark Built with 80+ Clinicians — OpenAI · 2026-09-24
- Paperena Benchmarks AI Scientists Across Full Research Cycles: Writing, Reviewing, Revising — yeewhye · 2026-09-24
- ACL'23 outstanding paper: discriminative LMs may generalize better than autoregressive models — ysu_nlp · 2026-09-24
- 10 agent reruns reached the right neighborhood, none reproduced the key observation — rohanpaul_ai · 2026-09-24
- GTSAM 4.3 ships with legged navigation, CUDA-accelerated factor graphs and continuous-time trajectory estimation — fdellaert · 2026-09-24