Continual learning today: harness engineering plus online RL from live interaction traces

willcb · x · 2026-09-19

willcb outlines how personal continual learning actually works today: it relies on careful harness engineering plus long-horizon RL for memory and multi-agent setups, with calibrated preference and reaction modeling as the key per-user learning target. Naive versions sort of work but are expensive and limited; the most effective real-world approach, he argues, is building task flywheels from live interaction traces and running online RL on them.

Related event: Online RL from Real Interaction Trajectories Gains Traction(2 posts)→

Original post →

More from coding & agent

coding & agent channel →