XM uses best-of-K training to improve fidelity and sample the data distribution
anshulkundaje · x · 2026-08-04
This thread explains the XM paper as a best-of-K style objective: the model is trained over many predictions, then optimized on the minimum loss, so it learns to better search the data distribution instead of collapsing toward the mean.
The author says this can make training faster and outputs higher fidelity, with free inference-time cost for the added expressivity. The thread also notes that best-of-K ideas are not new, but XM pushes the idea into a more general and useful form.
More from Research
- Pre-registered study finds a universal floor for hallucination detection, but no universal detector — k01234n · 2026-08-04
- Paper says a claimed “100% wrong” MATH set missed many valid answers — burny_tech · 2026-08-04
- Penn State talk explores how social science can help AI model moral judgment — soumitrashukla9 · 2026-08-04
- NeurIPS AC says reviewers ask for rebuttals but often ignore clarification requests — prajdabre · 2026-08-04
- Researchers teach a humanoid robot to learn from its own mistakes — imjustnewatai · 2026-08-04
- Qwen 3.8 Max reaches 42% on the hard INDUCTION benchmark, taking second place — DeryaTR_ · 2026-08-04