Scale AI launches RSI-Bench: $2,000 per accepted task plus co-authorship for contributors
dhruv2038 · x · 2026-08-18
Scale AI has launched RSI-Bench, an ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D — a prerequisite for recursive self-improvement. Each task ships with a starting environment, a compute budget, and a verifier suite measuring performance against a baseline. Beyond final outcomes, the benchmark evaluates reliability (distinguishing real gains from noise), efficiency (solving under limited resources), generality (solutions that extend beyond narrow fixes), and idea quality (extracting insights and synthesizing novel approaches).
The team is crowdsourcing expert-curated, long-horizon research tasks from the community. Contributors receive co-authorship on the RSI-Bench paper, $2,000 per accepted task for the initial 50 tasks, Modal compute credits, and access to the research community.
More from Research
- 10 YouTube Channels Worth Bookmarking for Learning Generative AI — goyalshaliniuk · 2026-08-18
- Berkeley Lab Unveils AI Model for Realistic Earthquake Simulation — Scobleizer · 2026-08-18
- Invented or discovered? Are neural networks a fundamental pattern of reality? — drabarca_ai · 2026-08-18
- KDD Cup Winners Unify Recommendation Systems, Team Built Winning Code with DeepSeek — 量子位 · 2026-08-18
- Spellcaster Uses 6-Agent Loop to Fix 'Unplayable' AI-Generated Games — 量子位 · 2026-08-18
- HumanCLAW open-sourced: all 9 SOTA VLMs fail embodied benchmark, best hits only 16.8% — liuziwei7 · 2026-08-18