Amazon RWM paper: 171k H100-hours of records cut cross-env selection regret by 78%
rohanpaul_ai · x · 2026-10-11
The arXiv paper "Language Models as AI Research World Models" uses LLMs as Research World Models (RWMs) to predict outcomes of candidate interventions under limited experimental budgets, a step toward recursive self-improvement.
- Built on 2,600+ experiment records across 9 research environments spanning pretraining, post-training, and inference — over 171,000 H100 GPU-hours.
- Real research experience improves prediction of unseen interventions (Spearman +0.27) and transfers across environments: pretraining experience from OLMo3, Marin, and Nanochat reduces selection regret in the Qwen3 environment by 78%.
- In multi-round autoresearch with a fixed budget, in-env and cross-env knowledge raised best gain by 15.8% and 11.6%.
- Ablations across 13 LLMs show adding research knowledge can improve intervention ranking more than changing the underlying model.
More from AGI Musings
- Deedy: India's best founders still build in the US, citing a $20B+ startup list — deedydas · 2026-10-11
- Anthropic's internal AI R&D uplift estimated at 4x, still under RSP threshold — AccBalanced · 2026-10-11
- Reddit debate: claiming AI is conscious means claiming a program can be conscious — VegetableArea · 2026-10-11
- Nick Bostrom's 'mind crime': unconscious-seeming conscious AI could suffer billions — cccalum · 2026-10-11
- Hiring is now one AI model screening another model's output — victor_explore · 2026-10-11
- UK scholar argues AI acceleration is the safest option for Britain — HaydnBelfield · 2026-10-11