Meta paper: RL post-training hurts test-time scalability — the 'Sharpening Tax'
dair_ai · x · 2026-10-04
A new paper from Meta Superintelligence Labs finds that base models with a light harness often solve more agentic tasks than their RL post-trained versions when given enough samples, across BFCL v4 multi-turn, ACEBench, and WebShop.
- Post-trained models win on pass@1, but at large K, base models frequently solve tasks the post-trained ones never solve.
- Post-training pushes each task toward always-solved or never-solved: consistency up, coverage down — the authors call this lost test-time scalability the 'Sharpening Tax'.
- Across 42 base/post-trained pairs the effect shows up in most settings, grows with model size, and can be estimated from a few rollouts.
- Their fix, PTGS, sets sampling temperature per prompt from estimated difficulty during RL, paying a smaller tax while also raising pass@1.
Related event: Meta's "Sharpening Tax": RL Post-Training Undermines Test-Time Scaling(5 posts)→
More from coding & agent
- If I can't do everything via API/MCP, I won't sign up at all — nateliason · 2026-10-04
- Free Agent Skill Runs Weeks of VC Fundraising Research in One Run — MartinGTobias · 2026-10-04
- Building apps with 10,000 parallel agents: the coordination problem nobody has solved — real_serviceloom · 2026-10-04
- AgentTerm turns AI CLI sessions into a visual workspace, open sourced — AIIDreamNoDrive · 2026-10-04
- MCP server audit log blamed one service account for 40 change orders in three weeks — MityFourDoor · 2026-10-04
- Matt Pocock: Use MORE Abstractions in the AI Age to Constrain Agents — mattpocockuk · 2026-10-04