vLLM precision gap prevents GRPO convergence

SergioPaniego · x · 2026-08-20

A technical finding shows that vLLM generating in bf16 while the trainer computes in fp32 stops GRPO from converging even on trivial tasks. The precision gap silently pushes tokens past PPO's clipping boundary, causing gradients to vanish. Matching precisions resulted in 5.8x better policy updates per step.

Original post →

More from Infra

Infra channel →