Why RL Environments Work Better in 2026: Greenblatt's Two Reasons

dejavucoder · x · 2026-09-12

Summarizing Ryan Greenblatt on the Dwarkesh podcast: RL environments work better in 2026 than 2024 because (1) we know what environments we need far better, and (2) 2026 models are much better at writing RL envs themselves, enabling human+AI high-quality and synthetic generation. The author suggests leveraging frontier computer-use models like Astra to speed up hard data annotation and mass-produce quality computer-use environments in knowledge work domains.

Original post →

More from Research

Research channel →