A sequence-level KL estimator can give the right gradient for async RL

ZacharyHuang12 · x · 2026-07-26

A technical thread argues that RL should be optimized at the sequence level, but most practical losses still collapse to token-level updates. The post highlights a paper that derives a sequence-level KL estimator with the correct gradient—which matters more than matching the exact KL value—and then extends the derivation to async RL.

Key point:

Original post →

More from coding & agent

coding & agent channel →