RSIBench-Data: Benchmarking Agents on Recursive Self-Improvement
imjustnewatai · x · 2026-07-29
This post provides the concrete data and links for the previously discussed agent automation paper (RSIBench-Data).
- Key Results: 14/24 settings improved beyond the first valid attempt. Of 23 searches continuing past their peak, 18 ended worse and 5 tied.
- Scope: One fixed target model, shared LoRA training, and six coding, terminal, science, and math benchmarks.
- Significance: The benchmark measures data-centric research capabilities rather than unrestricted self-rewriting.
Related event: RSIBench-Data: First Benchmark for Agent Recursive Self-Improvement(2 posts)→
More from Research
- HUG learns robot grasping from 1 million human handshapes and beats baselines by 34% — LerrelPinto · 2026-07-29
- Khora proposes a multi-agent world model that scales linearly instead of quadratically — siyuanhuang95 · 2026-07-29
- Sakana AI’s UnMaskFork uses multiple diffusion LMs to scale inference at test time — SakanaAILabs · 2026-07-29
- Nature Communications: Spatial Network Principles Underlying Neural Locomotion — plopesresearch · 2026-07-29
- Kevin Bryan Says Minimal RSI Could Arrive Around 2027 as AI R&D Productivity Rises 9% — imjustnewatai · 2026-07-29
- Feeding the Beast: Building JAX data pipelines with Grain — weskambale · 2026-07-29