RSIGym Launches as an Environment for AI to Improve AI, With RSI-Index Leaderboard
rohanpaul_ai · x · 2026-10-08
Evolvent AI released RSIGym and RSI-Index, a benchmark environment where autonomous research agents improve AI systems.
- Core idea: wraps five reusable research services — training, inference, rollout (data generation + judging), evaluation, and sandbox — into callable APIs, so agents spend their budget improving the target system instead of rebuilding infrastructure.
- Research loop: generate data → train → evaluate → analyze & revise, with a $500 budget per benchmark run, 6 research agents, and 5 benchmark domains.
- Three study modes: DATA (change training examples only), HARNESS (revise execution code with weights fixed), and JOINT (improve everything together, the main study).
- Scoring: agents submit a checkpoint + harness that goes through a fixed evaluation protocol for an official score, feeding the public RSI-Index leaderboard; the demo task is six agents improving Qwen3.5-35B-A3B-Base.
Paper, code, and website are all available.
Related event: Evolvent AI Launches RSIGym for Recursive Self-Improvement(3 posts)→
More from AGI Musings
- Frontier Models Decompiling Binaries Could Rescue Devices Bricked by Manufacturers — m4rkmc · 2026-10-08
- State of AI 2026: frontier narrows to three labs, inference costs drop 13x a year — Nathan Benaich (Air Street) · 2026-10-08
- Anthropic test of 52 devs: AI-assisted group scored 50% vs 67% without, error-finding worst — AlexTensor · 2026-10-08
- Dean Ball: The math-solving unreleased OpenAI model is still dumber than a corporation — deanwball · 2026-10-08
- IIM Bangalore replaces exams with building AI chatbots to answer test questions — CurieuxExplorer · 2026-10-08
- Ex-Anthropic security researcher talks superintelligence risk and alignment on Diary of a CEO — JeffLadish · 2026-10-08