NeurIPS 2026 paper: prompt interference explains why Pass@k optimization degrades Pass@1
lateinteraction · x · 2026-09-25
A NeurIPS 2026-accepted paper (arXiv:2602.21189) explains the recurring trade-off where inference-aware fine-tuning improves Pass@k but hurts Pass@1. The cause is 'prompt interference': Pass@k policy gradients implicitly upweight low-success prompts, and when these are negatively interfering, gradient conflicts rotate the update away from the Pass@1 direction. The authors give a theoretical characterization of when this occurs and validate it with LLM experiments on verifiable math reasoning tasks. The trade-off matters since Pass@1 remains a hard constraint due to latency, cost, and verifier coverage.
More from Models
- Meta launches Muse Code terminal coding agent, resets usage limits for DevDay — armand_ruiz · 2026-09-25
- One Week Left: GLM 5.3 Flash Tasks Cost $0.009 Until Sept 30, Then $0.09 — shensi · 2026-09-25
- How Jev, a new System 1 model, changes memory and context engineering — julianweisser · 2026-09-25
- ChatGPT confidently repeated a wrong iPhone price twice before user forced a check — InternationalFigure2 · 2026-09-25
- Side-by-side test puts Opus 5.5 against Astra, Sol, and Luna — Rasmic · 2026-09-25
- M5 Ultra 80-core tested with GLM-5.3-Flash: RAM is great, GPU is the bottleneck — dreamingwell · 2026-09-25