Amazon paper: LLM as research world model lifts experiment ranking correlation from 0.506 to 0.774
rohanpaul_ai · x · 2026-10-11
An Amazon team proposes using LLMs as Research World Models (RWMs) that predict whether a training intervention will help before any GPU time is spent, addressing the gap where AI research agents propose experiments faster than teams can afford to run them.
- Evaluation spans 2,653 real experiment records across 9 research environments (pretraining, post-training, inference), representing 171,000+ H100 GPU-hours.
- Adding past records from the same setup raised average ranking correlation with actual results from 0.506 to 0.774; on OLMo3-100M, low reasoning effort with records scored 0.892 vs 0.648 for max effort without records.
- Knowledge transfers across environments: pretraining experience from OLMo3, Marin, and Nanochat cuts selection regret in the Qwen3 environment by 78%.
- Under multi-round autoresearch with a fixed budget, in-env and cross-env knowledge improved best gain by 15.8% and 11.6%.
Practical takeaway: log every experiment, including failures, and feed those records to whatever model picks the next run.
More from AGI Musings
- Deedy: India's best founders still build in the US, citing a $20B+ startup list — deedydas · 2026-10-11
- Anthropic's internal AI R&D uplift estimated at 4x, still under RSP threshold — AccBalanced · 2026-10-11
- Reddit debate: claiming AI is conscious means claiming a program can be conscious — VegetableArea · 2026-10-11
- Nick Bostrom's 'mind crime': unconscious-seeming conscious AI could suffer billions — cccalum · 2026-10-11
- Hiring is now one AI model screening another model's output — victor_explore · 2026-10-11
- UK scholar argues AI acceleration is the safest option for Britain — HaydnBelfield · 2026-10-11