Evolvent AI releases RSIGym and RSI-Index: benchmarking AI self-improvement, Opus 5 leads at 0.4809
cihangxie · x · 2026-10-08
Evolvent AI argues RSI (recursive self-improvement) is a systems engineering problem, not just a model problem: progress depends on the research agent's environment—what resources it can call, what it can change, and how it runs experiments.
- RSIGym: an "Everything as a Service" environment exposing training, inference, evaluation, rollout, and sandbox as callable research services with unified auth and budget accounting, so agents focus on improving the target system.
- Three permission tracks: Data, Harness, and Joint (main study: data + training settings + harness together).
- RSI-Index: measures how well frontier agents jointly improve a target model's weights and harness. Across 6 research agents and 5 benchmark domains ($500 budget per run), Opus 5 leads at 0.4809.
Code, experiment configs, research trajectories, and eval logs are open-sourced.
More from coding & agent
- Demo is not delivery: what enterprise AI adoption actually takes — sujingshen · 2026-10-08
- DeLM's decentralized multi-agent system runs 2.49x faster, but MAS evals pick wildly different metrics — jyangballin · 2026-10-08
- Musk endorses Grok bot + Cursor cloud agents combo as a major productivity boost — elonmusk · 2026-10-08
- Dev: in the AI coding era, tabs-vs-spaces and style nitpicking never mattered — facontidavide · 2026-10-08
- RunningTab: environment-side task ledger consistently beats in-model tracking across 3 benchmarks and 3 LLMs — RexDouglass · 2026-10-08
- The top skill now: running overnight agents reliably without getting reward-hacked — zack_overflow · 2026-10-08