Why RL Environments Work Better in 2026: Greenblatt's Two Reasons
dejavucoder · x · 2026-09-12
Summarizing Ryan Greenblatt on the Dwarkesh podcast: RL environments work better in 2026 than 2024 because (1) we know what environments we need far better, and (2) 2026 models are much better at writing RL envs themselves, enabling human+AI high-quality and synthetic generation. The author suggests leveraging frontier computer-use models like Astra to speed up hard data annotation and mass-produce quality computer-use environments in knowledge work domains.
More from Research
- Year-Long Study of 1,000+ CharacterAI Users Finds AI Companionship Predicts Lower Well-Being — IanArawjo · 2026-09-12
- Fable 5.1 agent demo shows self-onboarding and lifelong learning at work — ysu_nlp · 2026-09-12
- Quadruped RL training in 1 minute on an M1 MacBook via mjbatch with 1024 parallel envs — philfung · 2026-09-12
- A tractable approach to pairwise interactions in ancestral sequence reconstruction — KevinKaichuang · 2026-09-12
- Generative Reward Models Fix Deceptive Autoformalization in Neurosymbolic Reasoning — CWRU · 2026-09-12
- AIRO launches automated catastrophic AI risk forecasts, matching top human forecasters — soumitrashukla9 · 2026-09-12