Scale AI's RSI Bench draws 600+ proposals, adds Schmidhuber as advisor
alexandr_wang · x · 2026-09-22
Alexandr Wang highlighted progress on Scale AI's RSI Bench, which measures whether AI agents can advance AI R&D through open-ended, iterative research: 600+ task proposals received, top contributors being onboarded to a public repo, and RSI pioneer Jürgen Schmidhuber joining as senior advisor.
The new verification blog explains the setup:
- Task design: each task gives an agent a fixed resource budget and a research environment with a reproducible baseline and a validation feedback loop for iterative experiments; final submissions are scored by a separate hidden evaluator under held-out conditions (unseen data, models, or environments).
- Verification pipeline: built on Harbor task formats, extending the Terminal-Bench verification workflow to iterative research; proposals are screened for under/over-specified instructions, verification that rewards only subsets of valid solutions, train/test splits rewarding memorization, hyperparameter sensitivity, and non-frontier baselines.
- Motivation: a benchmark only works if score increases reflect genuine capability gains, so proxy reward failures and exploitable evaluators must be ruled out.
More from Models
- Grok ships three frontier models in 9 weeks: 4.5, 4.6 and 4.7 back-to-back — XFreeze · 2026-09-22
- OpenAI reportedly rushing Codex Bot to rival Grok Bot, warns on recursive self-improvement — dotey · 2026-09-22
- Engineer autonomously trains a Jev-competitive model with an agent swarm for $3.1k in 20 hours — denisyarats · 2026-09-22
- Challenge: Track Your Daily Token Usage to Prove AI Companies Are Throttling Limits — tomchapin · 2026-09-22
- Day 1 with Grok 4.7: strict system-prompt adherence and visible gains over 4.5 in real coding work — elonmusk · 2026-09-22
- MoVA adds sparse value experts to attention for more capacity at no extra KV-cache cost — rupspace · 2026-09-22