100 rule-based verifiers and 300 tasks: team builds a working RL recipe for video models

DanielKhashabi · x · 2026-09-09

A team led by DengHokin proposes launching a "Verification Age" for video models, arguing they remain at the GPT-3 stage — trained by dumping millions of hours of video and hoping for magic — while coding models gained superhuman math/coding ability via synthetic RL environments with code/Lean verifiers.

Their recipe:

Built by 50 scientists over 6 months, the work argues rule-based verifiable rewards are the key to porting the RL paradigm from code to video.

Original post →

More from Multimodal

Multimodal channel →