SFT revival paper slammed for missing eval setup: single-shot sampling may make scores meaningless
hyunw_kim · x · 2026-10-05
A new paper claiming that adding sampling to the post-training stack lets SFT rival RL and OPSD—with better generalization and less forgetting—is drawing sharp criticism. A researcher points out the results section is a single paragraph, and the paper never specifies its eval setup: "single-shot accuracy" could mean one sample, zero-shot prompting, or one in-context example, with no temperature, decoding settings, or clarification of whether MMLU was scored by likelihood or generation. If the authors sampled once at nonzero temperature on 600-question test sets, that adds noise on top of seed variance—and for AMC, one sample per question would render the numbers nearly meaningless. The paper also oddly returns to experimenting with Qwen 2.5.
More from Models
- X rebuilds real-time reply moderation guide around Grok, dropping Perspective API — tetsuoai · 2026-10-05
- Alleged Opus 5.5 Generates an Eerie Face Laughing and Crying at Once — repligate · 2026-10-05
- Yacine: Opus hallucinates more, but Astra straight-up lies to me — yacineMTB · 2026-10-05
- Grok Bot defends no-roleplay design as users push for affectionate AI coworkers — repligate · 2026-10-05
- Gemini 4 Argon (High) tops Text Arena at 1525 pts, becomes top cost-efficient frontier model at $8/MToken — BLUECOW009 · 2026-10-05
- Nous Portal promo: 24 models 25-88% off, 9 models completely free — Teknium · 2026-10-05