Meta Quantifies the 'Sharpening Tax': RL Post-Training Trades pass@K Coverage for Accuracy

meta · hf · 2026-10-02

Meta's study confirms and refines the 'sharpening' hypothesis for LLM RL post-training: on agentic tasks, pre-trained models with a light inference harness often beat their post-trained counterparts in pass@K solution coverage despite far lower pass@1. Post-training pushes tasks toward always-solved or never-solved extremes, trading coverage for sampling efficiency. The team proposes Sharpening Tax, a diagnostic metric validated across 14 base/post-trained pairs from four families and three agentic benchmarks, estimable from a few rollouts. They also introduce posterior-tempered group sampling (PTGS), a plug-and-play Bayesian sampler that adapts temperature per prompt to estimated difficulty, paying a smaller tax while improving single-shot accuracy in two agentic RL environments.

Related event: Sharpening Tax: Post-Training Sharpens Old Skills Rather Than Teaching New Ones(2 posts)→

Original post →

More from Research

Research channel →