RSIBench-Data benchmarks whether agents can do data-centric research
rohanpaul_ai · x · 2026-07-29
RSIBench-Data tests whether agents can do data-centric research
The paper introduces RSIBench-Data, a controlled benchmark for recursive self-improvement focused on data-centric post-training research.
- It fixes the post-training stack so agents are evaluated on research decisions rather than optimization, serving, or systems work.
- Agents must iteratively revise training-data strategies for a fixed target model using feedback from official evaluation runs.
- The benchmark covers six tasks across software engineering, terminal use, scientific QA, and mathematics.
- Four frontier agents improved on the first valid attempt in 58.33% of settings, but performance was inconsistent.
- Among searches that continued past their best score, 78.26% ended with a worse final attempt, suggesting good candidate strategies often appear early or mid-run.
Related event: RSIBench-Data Tests Agents’ Data Research Skills(4 posts)→
More from coding & agent
- Five open-source tools promise 60%–95% fewer tokens for Claude and Codex agents — CodeByPoonam · 2026-07-29
- Codex Micro should be voice-first, with Git and PR controls replacing approve/reject — rudrank · 2026-07-29
- Agent teams should use deterministic checks first, then LLM judges, then humans — yoobinray · 2026-07-29
- Matt Pocock turns a wayfinding map into a full Cmd+K palette spec — mattpocockuk · 2026-07-29
- Codex can now handle trims, subtitles, masking and title cards in basic video edits — gabrielchua · 2026-07-29
- Anthropic warns Opus 5 may leak fake tool calls when thinking is off — heypearlai · 2026-07-29