Meta Paper: Post-Training Boosts pass@1 but Shrinks LLM Agents' pass@K Solution Coverage

iScienceLuvr · x · 2026-10-02

Meta researchers (Superintelligence Labs) introduce "Sharpening Tax," a diagnostic metric measuring how post-training reduces test-time scalability of LLM agents. Surprising finding: base models with a light inference harness often surpass post-trained counterparts in pass@K coverage despite lower pass@1, since post-training pushes per-task success toward 0% or 100%. They also propose PTGS (posterior-tempered group sampling), a plug-and-play Bayesian sampler that adapts temperature per prompt difficulty.

Related event: Meta's "Sharpening Tax": Post-Training Boosts First-Try Accuracy but Shrinks Solution Space(4 posts)→

Original post →

More from Research

Research channel →