OpenRSI-Index v0.1: benchmarking recursive self-improvement on 1K-GPU clusters

hhsun1 · x · 2026-09-24

OpenRSI released OpenRSI-Index v0.1, an open benchmark for whether AI agents can recursively improve themselves and extend the scientific frontier. It turns fully open-source projects into autoresearch environments on 1K-GPU production clusters, with agent trajectories lasting 60+ hours; building v0.1 took 100K+ H100-hours. The UPenn NLP group (hhsun1) noted it complements their work: ScienceAgentBench (ICLR 2025) for rigorous eval on expert-validated discovery tasks, AutoSDT (EMNLP 2025) for scaling scientific coding tasks from repos, and D3-Gym adding verifiable environments for data-driven discovery.

Related event: OpenRSI Launches Open Benchmark for Recursive Self-Improvement(2 posts)→

Original post →

More from coding & agent

coding & agent channel →