Sharpening Tax: RL post-training lifts agent pass@1 while shrinking pass@K below the base model

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li

cs.AI, cs.LG

2026-10-01

Post-training lifts agent pass@1 but shrinks pass@K: in 14 model pairs, base models plus a light harness overtake their RL twins by K=128. The loss is Sharpening Tax; PTGS cuts it.

What problem this solves

Does RL post-training add new capabilities, or does it just sharpen behaviors the base checkpoint already has? Almost all prior evidence comes from math and coding, the two domains least able to settle the question: pre-training corpora are saturated with math and code, and graders often check only the final answer, so a rollout with broken reasoning and a lucky answer still scores. pass@K gets inflated.

Agentic tasks are the real stress test. Multi-turn tool calling, environment feedback, and long trajectories are rare in pre-training data and widely assumed to be learned during post-training. A team from Meta Superintelligence Labs, UW-Madison, NYU, and Stanford ran 14 open base/post-trained checkpoint pairs on three agent benchmarks. Their conclusion: post-training mostly changes how reliably a model solves tasks, not which tasks it can solve.

Method

Three layers: measure the phenomenon, pin down the mechanism, then propose a metric and a fix.

Why adapt temperature per prompt? Global heating buys pass@K at the cost of pass@1, and group-contrast algorithms like GRPO only get a nonzero advantage when a group mixes successes and failures, which heating hard prompts makes more likely. PTGS leaves the RL update rule untouched.

Results

ComparisonMetricResult
WebShop, gemma-4-31B, K=128pass@128base+harness over 85%, post-trained 56%
Same, bucketed over 128 rolloutssolved-given-compute share87.6% down to 30.0%; never solved up from 12.4% to 44.0%
8 base models averaged, BFCLpass@32 with harness15.19 up to 49.13
42 model×benchmark casesTaxS(128) positive36/42
8-rollout tax predicting 32-rollout taxSpearman ρ0.85

The crossover budget shrinks with scale: on WebShop the 4B model needs more than 128 rollouts to overtake its post-trained twin, while 31B needs about 3. The smallest 3B backbones show negative tax at small budgets, so sharpening is a net win when compute is scarce.

PTGS was validated on Qwen2.5-7B-Instruct in Sokoban and FrozenLake, 200 steps, five seeds, PPO and GRPO:

Method (Sokoban)pass@1pass@128
base20.776.6
PPO46.555.0
PPO+PTGS61.169.7

In all four algorithm-by-environment combinations PTGS raised pass@1 and pass@128 at once and paid a smaller tax. Training logs show why: output entropy falls monotonically under fixed-temperature PPO, while PTGS keeps it high and repeatedly pushes it back up on prompts the policy keeps failing.

Why it matters

If your pipeline has a scalable verifier, a base model plus a harness plus repeated sampling can beat the instruct version on coverage. Scientific discovery and automated research, where one lucky rollout wins, are the natural fits. For single-call reliability, keep using the post-trained model.

The paper argues for reporting Sharpening Tax next to pass@1, and eight rollouts are enough to estimate it. PTGS touches only the sampler, so any group-sampling RL loop can adopt it. Keep proportions in mind, though: the diagnosis covers 42 cases, the fix one 7B model in two toy environments.

Limitations

The authors list four: open checkpoints do not disclose training data, so the analysis is observational rather than causal; RL, SFT, and distillation are not compared as separate tax sources; only text agents were tested; and the tax needs a binary success signal, which open-ended generation lacks.

Three more from this reading. RL post-trained is a loose label, and instruct checkpoints may be mostly SFT, which does not hurt the metric but weakens claims aimed at RL specifically. The PTGS experiments mix checkpoint selection rules (best-validation for GRPO, final for PPO), and Sokoban and FrozenLake sit far below BFCL or WebShop in complexity, so the fix has no evidence in realistic agent settings. The harness itself helps base models and hurts post-trained ones (their Table 4), so the exact crossover point partly depends on scaffolding choices.

Terms

Source

Related papers

All paper explainers