Meta Paper: Post-Training Boosts pass@1 but Shrinks LLM Agents' pass@K Solution Coverage
iScienceLuvr · x · 2026-10-02
Meta researchers (Superintelligence Labs) introduce "Sharpening Tax," a diagnostic metric measuring how post-training reduces test-time scalability of LLM agents. Surprising finding: base models with a light inference harness often surpass post-trained counterparts in pass@K coverage despite lower pass@1, since post-training pushes per-task success toward 0% or 100%. They also propose PTGS (posterior-tempered group sampling), a plug-and-play Bayesian sampler that adapts temperature per prompt difficulty.
More from Research
- OpenAI safety VP Lilian Weng shares her 7-step process for writing research blog posts — SinclairWang1 · 2026-10-02
- Arena.ai Launches HarnessTax: Quantifying How Much the Harness Matters for Coding Agents — solyarisoftware · 2026-10-02
- Berkeley paper: LLMs know your preference changed but still use the old one — rohanpaul_ai · 2026-10-02
- Frozen model, evolving harness: ModularRSI lifts Terminal-Bench 2.0 from 47.57 to 52.43 — jiqizhixin · 2026-10-02
- 30,000 paired QR-code illusions open-sourced with multi-decoder checks and robustness scores — 1roOt · 2026-10-02
- Google's Diffusion Controller: the 90% win rate and gray-box access refer to different setups — Crescitaly · 2026-10-02