OraRL: Efficient and Scalable RL for Video MLLMs

Yunheng Li · hf · 2026-08-26

OraRL improves RL post-training for video MLLMs by integrating oracle rollouts with decoupled advantage estimation and sign-balanced pruning. It achieves higher sample efficiency and scalability without chain-of-thought generation.

Original post →

More from Multimodal

Multimodal channel →