Yacine on RL for LLMs: everything behaves like training with a 400 gradient-step lag
cephaloform · x · 2026-09-24
Yacine shares a hands-on observation about RL on LLMs: nearly everything behaves as if the model is being trained with a lag of roughly 400 gradient steps. He also announces an upcoming interview with Databricks' mrdrozdov covering RL for knowledge agents, representation learning for retrieval, and the role of harnesses, soliciting questions on retrieval.
More from coding & agent
- Agentic Optimisation: what's fundamentally new for vision in the agent era — _krishna_murthy · 2026-09-24
- Dev uses Codex to turn a 3D-scanned firewood splitter into a native iOS game — OpenAIDevs · 2026-09-24
- Korean dev warns: non-coders using AI overlook everything invisible in software — algo_diver · 2026-09-24
- Anchor 3.0 launches as a runtime enforcement model for compliant AI agents — Saboo_Shubham_ · 2026-09-24
- Codex completes steps Claude refuses twice in one day, dev reports — ChanceKelch · 2026-09-24
- The bull case for assistant startups: MCP makes AI assistant apps interoperable — illscience · 2026-09-24