NeurIPS 2026 paper: prompt interference explains why Pass@k optimization degrades Pass@1

lateinteraction · x · 2026-09-25

A NeurIPS 2026-accepted paper (arXiv:2602.21189) explains the recurring trade-off where inference-aware fine-tuning improves Pass@k but hurts Pass@1. The cause is 'prompt interference': Pass@k policy gradients implicitly upweight low-success prompts, and when these are negatively interfering, gradient conflicts rotate the update away from the Pass@1 direction. The authors give a theoretical characterization of when this occurs and validate it with LLM experiments on verifiable math reasoning tasks. The trade-off matters since Pass@1 remains a hard constraint due to latency, cost, and verifier coverage.

Original post →

More from Models

Models channel →