The weakest of 13 LMs with experiment logs outranks the strongest without

Language Models as AI Research World Models

Zijun Wang, Zewen Liu, Minhua Lin, Zhaotian Weng, Zhan Shi, Bing He, Yisi Sang, Dakuo Wang, Benoit Dumoulin, Wei Jin, Yuyin Zhou, Cihang Xie, Hanqing Lu

cs.CL

2026-10-09

Amazon-led RWM study: 2,653 real runs lift in-env Spearman from 0.506 to 0.774, and OLMo3/Marin/Nanochat logs cut Qwen3 selection regret by 78%.

What problem this solves

Research agents can propose interventions faster than those interventions can be run. A learning-rate schedule, data mixture, optimizer, or architecture edit can help under one base model and budget, and do nothing or hurt under another. When the budget is fixed, ranking candidates changes the pace of progress more than generating extra ideas.

World models choose actions by predicting their consequences. Here the action is an executable research edit, and the consequence is gain against a fixed reference recipe. A Research World Model (RWM) is one language model that reads an environment description plus past runs and estimates the gain of an intervention that has not been executed. The sharper question is whether real records from environment A help predict unseen interventions in environment B.

Method

The weights stay frozen. The input is the environment (reference model, data, budget, eval protocol), the intervention (a code diff or config override), and a history of records. The output is the conditional median gain. For bits per byte (BPB, lower is better), gain is reference minus candidate. For accuracy and throughput, gain is candidate minus reference. Positive means an improvement.

Three history conditions share the backbone and the inference settings. W0 receives no records. In-env history contains only past runs from the target environment. Cross-env history contains other environments only: the target's own outcomes stay out of context, and near-duplicates of the queried intervention are removed from the source records. Records are serialized into the prompt. The model returns a median and may not call tools or search. The default backbone is Claude Opus 5 with a 1M-token context and HIGH reasoning effort.

The baselines are k-nearest-neighbor and ridge regression, fit on the same records. Vocabularies, TF-IDF weights, and hyperparameters use only the permitted history. Selection is simulated by executing 3 of 16 candidates, about a fifth of each pool, averaged over 200 shared random partitions.

Results

The authors' own runs supply 2,653 records from nine environments and 171,862 H100 GPU-hours. In-env prediction covers five environments. Cross-env prediction stays inside five autoregressive pretraining environments and uses 244 source records each time.

In-env records lift mean Spearman rank correlation from 0.506 to 0.774. On both MAE and Spearman, the record-conditioned RWM beats kNN and ridge in all five environments.

EnvironmentSpearman, no recordsSpearman, with recordsRegret@3, no recordsRegret@3, with records
OLMo3-100M0.620.902.80e-3 BPB7.00e-4
Diffusion pretraining0.460.783.14e-3 BPB4.70e-4
Math distillation0.430.700.48 pp0.29 pp
Code RL0.480.622.20 pp0.83 pp
Inference optimization0.540.883.60 tok/s2.53 tok/s

The equal-weight mean relative drop in Regret@3, the gain lost by missing the best item when executing 3 of 16, is 58.2%. NDCG@3, which scores only positive gains, rises from 0.599 to 0.785. Ranking favors the RWM in every environment. Selection regret does not: on inference optimization, kNN Regret@3 is 1.62 tok/s and the record-conditioned RWM is 2.53.

Cross-env macro Spearman moves from 0.628 to 0.729, and macro Regret@3 falls 52.2% relative to W0. On Qwen3 the drop is from 9.40e-4 to 2.10e-4 BPB, 78%, using only OLMo3, Marin, and Nanochat. kNN falls below W0 on all five transfer targets, with macro Spearman 0.36. Nanochat already uses Muon. Source analogues of a Muon variant were mostly helpful, so kNN predicted +0.018 BPB. The measured gain was −0.036, and W0 predicted −0.035. Similar edits point the wrong way once the recipe changes.

Multi-round Autoresearch on OLMo3-100M is repeated five times, eight rounds each. Every round adds 16 candidates and each method executes 3, for 24 selections in total. Against no initial records, in-env history raises mean final best gain by 15.8%, and cross-env history raises it by 11.6%. Normalized area under the best-so-far curve, with an Arrival Oracle that already knows the measured gains set to 100%, is 85.26, 98.26, and 94.34 for W0, in-env, and cross-env. kNN and ridge given the same cross-env records reach 86.16 and 78.06.

Across 13 backbones, including Claude Opus 5, Grok-4.6, the GPT-5.6 series, Kimi-K2.5, and DeepSeek-V3.2, Spearman without records spans −0.362 to 0.621 and with records spans 0.638 to 0.896. The ranges do not overlap. Moving Claude Opus 5 from LOW to MAX reasoning lifts Spearman only from 0.609 to 0.648. LOW plus records reaches 0.892.

Scoring 16 candidates once each costs about $6.3 on Claude Opus 5. The median experiment across the nine environments costs 42.3 H100 hours, about $127 at $3 per GPU-hour. One scoring pass is about 5% of one representative run.

Why it matters

The scarce resource is the execution slot. Logs that were already paid for should keep the failures too, because the next choice needs the whole gain distribution.

On this panel, adding records moves rankings more than swapping the backbone, and more than maxing out reasoning effort. The gain comes from in-context learning. A trained research world model is still future work, and the step is incremental. Teams that already hold large internal experiment archives get more from wiring those logs into selection.

Limitations

Cross-environment evidence is only autoregressive pretraining. Diffusion, math distillation, code RL, and inference optimization are in-env only. Inside an environment every run starts from the same reference recipe, so a mid-project change of base model is untested.

Point estimates of cross-env Spearman beat W0 on all five targets, but paired bootstrap intervals include zero for two of them: OLMo3-190M is +0.06 [−0.07, +0.18], and Marin is +0.08 [−0.05, +0.24]. Marin selection moves the wrong way as well. Regret@3 goes from 7.40e-4 to 7.90e-4, and NDCG@3 from 0.74 to 0.70.

Code RL has 34 history records against 115 test items. Reference reruns of autoregressive pretraining show a standard deviation of about 0.003 to 0.006 mean BPB, and intervention-level repeat variance is not measured. In-env MAE on OLMo3-100M falls from 2.93e-2 to 2.03e-2 BPB, still several times that reference noise. Ranking and selection are the firmer results. Absolute gain forecasts remain loose.

Predictions go through a commercial API and are not bitwise reproducible. The multi-round protocol uses a closed pool: 231 candidates are sampled in advance and released over eight rounds by mechanism complexity, and the methods only choose. A better forecast still has to be judged against the full bill of data collection, inference, and execution. Transfer into environment types with no related records, and whether training a dedicated RWM on these logs beats in-context use, are both untested.

Terms

Source

What people are saying

Related papers

All paper explainers