Diverse Approach Sampling Beats IID Self-Training on Hard Problems
PMinervini · x · 2026-10-04
Researchers at the University of Edinburgh and CIFAR propose changing how self-training data is sampled: instead of independently sampling answers and keeping the correct ones (RFT), have the model first propose several distinct approaches, then solve the problem once per approach.
- Problem: On frontier problems (solved at most 2 of 64 times), IID samples tend to be mode-collapsed and repeat one strategy, so fine-tuning barely shifts behavior. With 4 samples per training problem, Qwen3-4B-Instruct's held-out pass@64 stays at 5.2; even 64 samples at temperature 1.5 only reach 7.9.
- Method: Two ways to elicit approaches — GROOT (the model writes a decision tree of approaches and n root-to-leaf paths are taken) and Verbalized Sampling (a list of approaches).
- Result: Approach-diverse sampling yields better pass@k, test-time scaling and RL initialization.
More from Research
- Ontology Pipeline: a semantic infrastructure framework for the AI era — adnan_hashmi · 2026-10-05
- How 12 Napoleonic-era telegrams were cracked via an alphabetical codebook — apples_jimmy · 2026-10-05
- The Silhouette Fallacy: why copying neuron geometry isn't inheriting its computation — neurovium · 2026-10-05
- dronegeo: simulation study compares racing drone geometries on lap times — DominiqueCAPaul · 2026-10-05
- Running Qwen3.5 9B/27B INT4 on cheap ex-mining FPGA boards — I_am_purrfect · 2026-10-05
- Researchers Debate RL Generalization Limits: If It Generalized Well, Labs Wouldn't Need to Build Envs — brianryhuang · 2026-10-05