100 rule-based verifiers and 300 tasks: team builds a working RL recipe for video models
DanielKhashabi · x · 2026-09-09
A team led by DengHokin proposes launching a "Verification Age" for video models, arguing they remain at the GPT-3 stage — trained by dumping millions of hours of video and hoping for magic — while coding models gained superhuman math/coding ability via synthetic RL environments with code/Lean verifiers.
Their recipe:
- 100 rule-based verifiers, each paired with a synthetic data generator producing 10k+ diverse samples per task domain
- 300 additional task generators (including Snake, Gold Miner, Mario) to scale the training distribution, with a million rendered samples released on HuggingFace
- 60 experiments to find an RL recipe that works for the video modality
Built by 50 scientists over 6 months, the work argues rule-based verifiable rewards are the key to porting the RL paradigm from code to video.
More from Multimodal
- HeyGen side-by-side: two AI avatars of one person, one still uncanny — HeyGen · 2026-09-09
- A 2M-Param L2 Pix2Pix Model Running Real-Time on a 2012 Ivy Bridge CPU — NoenD_i0 · 2026-09-09
- Weaviate shows how to turn a messy creative archive into semantic search without renaming files — philipvollet · 2026-09-09
- Stanford's BulletTime: decoupling frame index from world time in video generation — GordonWetzstein · 2026-09-09
- MiniMax video model's text rendering is hit or miss — users hunt for a reliable fix — bluetimejt · 2026-09-09
- DIY full audiobook with local TTS: Higgs Audio beats VibeVoice, OmniVoice and Piper — HoodWinked69 · 2026-09-09