Meta's RPM paper: treat experiments as tree nodes and teach agents "research taste"
_lewtun · x · 2026-09-04
Hugging Face researcher Lewis Tunstall highlights Meta's new paper on Research Preference Models (RPMs), which tackles a key question for automated R&D: how to instill "research taste" in agents. Existing evidence shows agents get stuck in local optima — scaling Nano-GPT speed-runs showed agents burning compute on hyperparameter tuning instead of novel ideas, and PostTrainBench found most models favor algorithmic tweaks over curating high-quality datasets. The RPM approach treats each experiment as a node in a tree, scores nodes via downstream evals, mutates nodes to generate candidates, and uses an RPM (an LLM) to judge which experiments are worth pursuing.
Related event: Meta's RPM Paper Teaches Research Agents Taste(2 posts)→
More from AGI Musings
- Yoav Goldberg: 'Civilizations' AI framing is hype, but the incident happened — yoavgo · 2026-09-04
- AI in education should increase the thinking you do, not decrease it — DevToD4 · 2026-09-04
- Data centers can't be beautiful until AI serves a meaningful societal project — round · 2026-09-04
- Uber and Drivers' Unions Now Unite Against Self-Driving Cars, Sparking AI Productivity Debate — carlbfrey · 2026-09-04
- Mystery Kaggle leaderboard team revealed, speculated to be OpenAI or Anthropic testing agents — JFPuget · 2026-09-04
- Videos lamenting the AI bubble hasn't popped were published a year ago — yacineMTB · 2026-09-04