Online RL via task flywheels from live interaction traces is what actually works today
willcb · x · 2026-09-19
Responding to a thread on whether models will robustly solve complex tasks and become economically viable, the author argues that naive approaches can roughly work but are expensive, fiddly and limited in scope.
The version that works best today for real workloads is carefully constructed task flywheels built from live interaction traces, fed into online RL.
Related event: Online RL from Real Interaction Trajectories Gains Traction(2 posts)→
More from coding & agent
- Mnemos engine builds synaptic-like agent memory via resonant engrams — RileyRalmuto · 2026-09-19
- Codex's built-in browser now supports extensions and cookie import — vista8 · 2026-09-19
- Jev's structured-output model impresses, Yegge argues fences beat sandboxes for agents — njyx · 2026-09-19
- Archify, a 67k-star open-source skill, turns one-line system descriptions into interactive diagrams — alex_verem · 2026-09-19
- Generating a 40s motion graphic video with GLM writing code, no video model — 9r4n4y · 2026-09-19
- UnCorreoTemporal MCP server lets AI agents grab OTP codes and verification links with zero human help — modelcontextprotocol · 2026-09-19