Tobias Lee: RL for Verifiable Tasks, MOPD for Open Domains
_AndrewZhao · x · 2026-09-17
In a reply to AndrewZhao, Tobias Lee lays out a terse take on training directions: use reinforcement learning for verifiable tasks, and MOPD for open domains — highlighting the split between verifiable and non-verifiable problems in current post-training approaches.
More from Research
- Harvard study suggests smartphones could predict suicide risk with striking accuracy — pshrink · 2026-09-17
- Study of 7 models across Claude Code, Codex, Pi: harness barely affects success but swings cost — DavideCrapis · 2026-09-17
- Neuroevolved value function solves Tetris at world-record speed in your browser — NathanWilbanks_ · 2026-09-17
- Polymarket atomicity gap: 1.8M reverted trades expose ghost-filled order attack — chaumian · 2026-09-17
- NeurIPS 2026 opens financial aid and volunteer applications, due Oct 6 — NeurIPSConf · 2026-09-17
- Classifying 1,018 AI Papers for $4: A Two-Model Pipeline at 256ms Median Latency — nutlope · 2026-09-17