Continual learning today: harness engineering plus online RL from live interaction traces
willcb · x · 2026-09-19
willcb outlines how personal continual learning actually works today: it relies on careful harness engineering plus long-horizon RL for memory and multi-agent setups, with calibrated preference and reaction modeling as the key per-user learning target. Naive versions sort of work but are expensive and limited; the most effective real-world approach, he argues, is building task flywheels from live interaction traces and running online RL on them.
Related event: Online RL from Real Interaction Trajectories Gains Traction(2 posts)→
More from coding & agent
- Docling, the Document Parser for Gen AI, Tops 66K GitHub Stars — docling-project · 2026-09-19
- Codex-X Adds a Visual Management Panel for OpenAI Codex Desktop and CLI — yynxxxxx · 2026-09-19
- 'Letting Agents Rip' on Your Codebase: The Coding Agent Meme Everyone Gets — charles_irl · 2026-09-19
- Claude Code adds AGENTS.md support in 2.1.277, falling back when CLAUDE.md is absent — eyishazyer · 2026-09-19
- Open-source Jev-cu makes Codex computer use faster by passing text, not screenshots — alexcovo_eth · 2026-09-19
- Jev read 384 news stories for $0.19 while Claude Opus 5 managed 4 for $0.77, claims demo — airesearch12 · 2026-09-19