OpenRSI-Index v0.1: benchmarking recursive self-improvement on 1K-GPU clusters
hhsun1 · x · 2026-09-24
OpenRSI released OpenRSI-Index v0.1, an open benchmark for whether AI agents can recursively improve themselves and extend the scientific frontier. It turns fully open-source projects into autoresearch environments on 1K-GPU production clusters, with agent trajectories lasting 60+ hours; building v0.1 took 100K+ H100-hours. The UPenn NLP group (hhsun1) noted it complements their work: ScienceAgentBench (ICLR 2025) for rigorous eval on expert-validated discovery tasks, AutoSDT (EMNLP 2025) for scaling scientific coding tasks from repos, and D3-Gym adding verifiable environments for data-driven discovery.
Related event: OpenRSI Launches Open Benchmark for Recursive Self-Improvement(2 posts)→
More from coding & agent
- Two Codex agents coded, tested and submitted iOS & Android apps to stores in parallel — burkov · 2026-09-24
- You misunderstand AGENTS.md: how coding tools read project instructions — dotey · 2026-09-24
- Forced fresh-worker swap at 60k tokens: 10/12 Terminal-Bench tasks still pass — key_of_door · 2026-09-24
- Solo Dev Open-Sources Tapioca, a Go Terminal Coding Agent That Runs Fully Local — Practical_Witness_95 · 2026-09-24
- Anthropic launches Claude Marketplace with 2,000+ connectors and third-party agents — claudeai · 2026-09-24
- Nyx Terminal launches for Mac: a terminal built for running many coding agents at once — henrymodis · 2026-09-24