Hill Sampling Beats Repeated Sampling and Test-Time Training by Conditioning on Best Program

burny_tech · x · 2026-09-29

A new arXiv paper (Jacob Beck, Philip V. Ogren, Ari Kobren) proposes Hill Sampling as a simpler, better alternative to repeated sampling, evolution, and test-time weight training for test-time scaling.

Core idea: condition a frozen LLM on the best program it has found so far, iterating via a simple prompt-conditioned hill-climbing loop.

Result: this plain approach matches or beats complex test-time weight training and evolutionary search. Practitioners spending massive compute on evolutionary loops or fine-tuning for code and discovery may get equal or better results with the simple loop instead.

Original post →

More from Research

Research channel →