Calibrated Importance Sampling Fixes Training-Inference Mismatch in LLM RLVR
Tianrun Yu · hf · 2026-09-29
This work studies training-inference mismatch in RLVR: rollouts are sampled by an inference engine while gradients are computed by a training engine, which assign different probabilities to the same tokens, biasing policy updates.
- Method: calibrated importance sampling (CIS), motivated by an empirically supported logit-displacement characterization expressing mismatch as additive displacement in log-odds, approximately invariant to token confidence. This yields confidence-aware truncation: large positive displacements truncated at a single constant threshold, mapping to an importance-ratio cap that tightens as confidence rises.
- Theory: CIS replaces the unbounded second moment governing exact importance sampling error with a constant-bounded term, at the cost of a controllable bias.
- Results: CIS achieves the best five-benchmark average on all three MoE models tested across five math reasoning benchmarks; truncation bias on low-confidence tokens is lower than truncated importance sampling, while upward clipping of small importance weights hurts held-out accuracy.
More from Research
- Ten Claude agents prove Thomson problem (N=7) with a 17,895-line Lean proof in 15 hours — aran_nayebi · 2026-09-29
- Undergrad Gets Personal ML Research Into NeurIPS Poster, Asks About the Vibe — XxCotHGxX · 2026-09-29
- Are We Optimizing the Wrong Metric? Community Debates ML Benchmarks vs Real Value — Physical_Tea9389 · 2026-09-29
- UCLA lands $25M federal grant to lead national evaluation of AI tools for Alzheimer's — chrismattmann · 2026-09-29
- Researchers flag LLM checking limits: validation is fluent, not formally verified — anshulkundaje · 2026-09-29
- Foresight hosts SF conference on AI-first science with DeepMind, MIT speakers — juanbenet · 2026-09-29