UCL's Large Discovery Models hit SOTA, 2.4x LLM reflection on experiment search

jiqizhixin · x · 2026-09-08

Jun Wang's group at UCL presents Large Discovery Models (LDM), turning model-based open-ended search into an empirical loop: given experimental data and a foundation-model context, it builds a Bayesian reward signal over the search space, scoring each candidate by performance or uncertainty reduction. A fast loop iterates in real time from experimental data; a slow loop distills the reward signal into the foundation model via post-training. On H100, LDM achieves 2.4x the BPB reduction of pure LLM reflection under identical conditions; on B200 it actively widens the search space, driving BPB down to 0.902291 — world #1 on the Auto-R leaderboard.

Original post →

More from Research

Research channel →