Tobias Lee: RL for Verifiable Tasks, MOPD for Open Domains

_AndrewZhao · x · 2026-09-17

In a reply to AndrewZhao, Tobias Lee lays out a terse take on training directions: use reinforcement learning for verifiable tasks, and MOPD for open domains — highlighting the split between verifiable and non-verifiable problems in current post-training approaches.

Original post →

More from Research

Research channel →