SFT revival paper slammed for missing eval setup: single-shot sampling may make scores meaningless

hyunw_kim · x · 2026-10-05

A new paper claiming that adding sampling to the post-training stack lets SFT rival RL and OPSD—with better generalization and less forgetting—is drawing sharp criticism. A researcher points out the results section is a single paragraph, and the paper never specifies its eval setup: "single-shot accuracy" could mean one sample, zero-shot prompting, or one in-context example, with no temperature, decoding settings, or clarification of whether MMLU was scored by likelihood or generation. If the authors sampled once at nonzero temperature on 600-question test sets, that adds noise on top of seed variance—and for AMC, one sample per question would render the numbers nearly meaningless. The paper also oddly returns to experimenting with Qwen 2.5.

Original post →

More from Models

Models channel →