Meta paper: RL post-training hurts test-time scalability — the 'Sharpening Tax'

dair_ai · x · 2026-10-04

A new paper from Meta Superintelligence Labs finds that base models with a light harness often solve more agentic tasks than their RL post-trained versions when given enough samples, across BFCL v4 multi-turn, ACEBench, and WebShop.

Related event: Meta's "Sharpening Tax": RL Post-Training Undermines Test-Time Scaling(5 posts)→

Original post →

More from coding & agent

coding & agent channel →