Claude Opus 5 nearly triples Qwen's SWE-bench score in open-source RSIGym auto-research env
rohanpaul_ai · x · 2026-10-09
EvolventAI released RSIGym, an open-source environment that lets AI research agents retrain a model and rewrite its evaluation harness under a single budget, testing whether frontier agents can act as automated AI researchers.
Key points:
- Everything-as-a-Service design: agents run from CPU-only containers and call remote services for LoRA fine-tuning, serving, benchmarking and sandboxes, all charged against a per-run budget.
- In the main test, 6 frontier agents started from Qwen3.5-35B-A3B-Base with a minimal harness and $500 of services per benchmark; Claude Opus 5 nearly tripled Qwen's SWE-bench Verified score and led RSI-Index at 0.4809.
- RSI-Index measures how well agents jointly improve a target model's weights and harness, across 6 agents and 5 benchmarks.
- Practical takeaway: read failure logs and fix the harness first — heavier training scored lower in 8 of 10 small tests.
- Author's thesis: recursive self-improvement is a systems engineering problem, not just a model problem; progress hinges on what the research environment lets agents call and change.
Related event: Evolvent AI Open-Sources RSIGym for Self-Improving AI Research Agents(5 posts)→
More from coding & agent
- Swapping harness lifts GPT-5.6 repo migration from 6.5% to 31%, paper finds — omarsar0 · 2026-10-09
- He replaced a 17-step agent prompt checklist with a 1976 Unix Makefile — 140+ tickets without drift — RandalSchwartz · 2026-10-09
- Developer builds wildfire analysis tool with Opus 5.5: fire progression and hot spot tracking — natesiggard · 2026-10-09
- Subagent-as-a-service: betting on 'expert' agents like Harvey as tool calls for vertical AI — nbaschez · 2026-10-09
- WebMCP Debate: LLM Agents Skip Ads, Putting Ad-Based Business Models at Risk — ConnectRub4818 · 2026-10-09
- At first Grok Bot Meetup, audience-voted idea becomes a 139-marker map app in ~10 minutes — pswider · 2026-10-09