Beyond RL: Why Environments Are the Highest-Lever Way to Build AI Agents
ben_burtenshaw · x · 2026-08-13
The author argues that the classic ML MATTER cycle (model, annotate, train, test, evaluate, revise) needs an upgrade for modern agentic workflows.
He believes that environments (and to some extent datasets) are now the primary way capabilities get into models. Currently, the industry lumps environments together with RL and post-training, which is a limitation. Going forward, environments should be treated as the core mechanism to discover, measure, and improve model capabilities—whether in the weights, the harness, or the prompt—serving as the highest-leverage way for builders to customize AI.
More from coding & agent
- Conclave Personal: open-source multi-agent verification tool assigns roles to LLMs for self-correction — HospitalSlight7930 · 2026-08-13
- Can LLMs Self-Correct? Conclave Uses Multi-Agent Roles to Verify Outputs — ProposalIntrepid8476 · 2026-08-13
- Building an Always-On Personal Agent with Obsidian and Hermes — hugobowne · 2026-08-13
- SkillZip: Contract-Preserving Graph Compression for Agent Skill Libraries — Xingyu Tan · 2026-08-13
- Human Edge in the Agentic Era: Delegate Pattern Matching, Keep Human Expertise — blaizedsouza · 2026-08-13
- Claude 3.5 Sonnet Excels at Frontend Design with Screenshot Iteration — iannuttall · 2026-08-13