Hill Sampling Beats Repeated Sampling and Test-Time Training by Conditioning on Best Program
burny_tech · x · 2026-09-29
A new arXiv paper (Jacob Beck, Philip V. Ogren, Ari Kobren) proposes Hill Sampling as a simpler, better alternative to repeated sampling, evolution, and test-time weight training for test-time scaling.
Core idea: condition a frozen LLM on the best program it has found so far, iterating via a simple prompt-conditioned hill-climbing loop.
Result: this plain approach matches or beats complex test-time weight training and evolutionary search. Practitioners spending massive compute on evolutionary loops or fine-tuning for code and discovery may get equal or better results with the simple loop instead.
More from Research
- EAPO: entropy-guided credit assignment for RLVR improves exploration in LLM reasoning — coallaoh · 2026-09-29
- Atlases Are Already Inside: Recovering Population Templates by Making Diffusion Models Collapse — kwangmoo_yi · 2026-09-29
- Duplex-MPE benchmarks multi-party full-duplex speech: MiniCPM-o 4.5 leads on three of four scores — PKU · 2026-09-29
- PReCache: Training-Free KV Cache Sharing Gives Multi-LoRA Agents up to 3.1x TTFT Speedup — SNU-VLSI · 2026-09-29
- REALM generates reactive listener facial motion, deployed on an Ameca humanoid robot — MacquarieUni · 2026-09-29
- KAIST's FlyBy teaches small models when to query stronger ones, beating Qwen3-14B at 2.7x lower cost — kaist-ai · 2026-09-29