Meta Quantifies the 'Sharpening Tax': RL Post-Training Trades pass@K Coverage for Accuracy
meta · hf · 2026-10-02
Meta's study confirms and refines the 'sharpening' hypothesis for LLM RL post-training: on agentic tasks, pre-trained models with a light inference harness often beat their post-trained counterparts in pass@K solution coverage despite far lower pass@1. Post-training pushes tasks toward always-solved or never-solved extremes, trading coverage for sampling efficiency. The team proposes Sharpening Tax, a diagnostic metric validated across 14 base/post-trained pairs from four families and three agentic benchmarks, estimable from a few rollouts. They also introduce posterior-tempered group sampling (PTGS), a plug-and-play Bayesian sampler that adapts temperature per prompt to estimated difficulty, paying a smaller tax while improving single-shot accuracy in two agentic RL environments.
More from Research
- Nature cover: 6.3M nuclei sequenced to build largest human prefrontal cortex cell atlas — jiqizhixin · 2026-10-02
- Detect LLM hallucinations in 1.3µs on CPU — but 120B models hallucinate with unanimous false certainty — More_Slide5739 · 2026-10-02
- 18-Year-Old Wins $100K Top ISEF Prize for MCMC Method Sampling Origami Motions — burny_tech · 2026-10-02
- Why 'Weird' Neurons Matter: Mixed Selectivity in Cortex Through the Lens of Cover's Theorem — burny_tech · 2026-10-02
- Adaptive effort is likely the killer app for dynamic looped transformers — willcb · 2026-10-02
- HuST Lab's Multimodal Flow: Fully Continuous Unified Language-Vision Generation — hustvl · 2026-10-02