Yacine on RL for LLMs: everything behaves like training with a 400 gradient-step lag

cephaloform · x · 2026-09-24

Yacine shares a hands-on observation about RL on LLMs: nearly everything behaves as if the model is being trained with a lag of roughly 400 gradient steps. He also announces an upcoming interview with Databricks' mrdrozdov covering RL for knowledge agents, representation learning for retrieval, and the role of harnesses, soliciting questions on retrieval.

Original post →

More from coding & agent

coding & agent channel →