Calibrated Importance Sampling Fixes Training-Inference Mismatch in LLM RLVR

Tianrun Yu · hf · 2026-09-29

This work studies training-inference mismatch in RLVR: rollouts are sampled by an inference engine while gradients are computed by a training engine, which assign different probabilities to the same tokens, biasing policy updates.

Original post →

More from Research

Research channel →