Andrew Zhao: MOPD is an org advantage — expert teams run their own RL, deliver experts

_AndrewZhao · x · 2026-09-17

In a thread with Roy Xie, Andrew Zhao argues MOPD is more of an organizational advantage than a performance one: code and reasoning teams can each run their own RL (critic or not, choosing their own hyperparameters and algorithms) to produce their best expert, then hand off that expert and data as a deliverable to model integration, instead of the integration team doing all the RL.

The integration team only needs to figure out the KL matching setup, a supervised learning objective that is far easier than RLing everything. Xie remains curious whether MOPD is actually needed at all.

Original post →

More from Companies & People

Companies & People channel →