Online RL via task flywheels from live interaction traces is what actually works today

willcb · x · 2026-09-19

Responding to a thread on whether models will robustly solve complex tasks and become economically viable, the author argues that naive approaches can roughly work but are expensive, fiddly and limited in scope.

The version that works best today for real workloads is carefully constructed task flywheels built from live interaction traces, fed into online RL.

Related event: Online RL from Real Interaction Trajectories Gains Traction(2 posts)→

Original post →

More from coding & agent

coding & agent channel →