Andrew Zhao: MOPD is an org advantage — expert teams run their own RL, deliver experts
_AndrewZhao · x · 2026-09-17
In a thread with Roy Xie, Andrew Zhao argues MOPD is more of an organizational advantage than a performance one: code and reasoning teams can each run their own RL (critic or not, choosing their own hyperparameters and algorithms) to produce their best expert, then hand off that expert and data as a deliverable to model integration, instead of the integration team doing all the RL.
The integration team only needs to figure out the KL matching setup, a supervised learning objective that is far easier than RLing everything. Xie remains curious whether MOPD is actually needed at all.
More from Companies & People
- Slack launches Slack Code for agentic team coding with 11 agents incl. Replit — amasad · 2026-09-17
- LangChain CEO: our gateway shipped a year too late, sub-agents rarely work as expected — Hacubu · 2026-09-17
- Anthropic merges Claude Chat and Cowork, launches ClaudeDocs and ClaudeSlides — 量子位 · 2026-09-17
- Ex-Microsoft engineer recalls Windows 8 meetings: note app idea met with abuse, then he quit — lmoroney · 2026-09-17
- MSK Partners With OpenEvidence, Binging OncoKB Precision Oncology Data to Doctors — benxneo · 2026-09-17
- Manus Launches Global "Learn AI with Manus" Partner Program, Starting with Singapore — parker_lyman · 2026-09-17