Meta's RPM paper: treat experiments as tree nodes and teach agents "research taste"

_lewtun · x · 2026-09-04

Hugging Face researcher Lewis Tunstall highlights Meta's new paper on Research Preference Models (RPMs), which tackles a key question for automated R&D: how to instill "research taste" in agents. Existing evidence shows agents get stuck in local optima — scaling Nano-GPT speed-runs showed agents burning compute on hyperparameter tuning instead of novel ideas, and PostTrainBench found most models favor algorithmic tweaks over curating high-quality datasets. The RPM approach treats each experiment as a node in a tree, scores nodes via downstream evals, mutates nodes to generate candidates, and uses an RPM (an LLM) to judge which experiments are worth pursuing.

Related event: Meta's RPM Paper Teaches Research Agents Taste(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →