VBVR-Pro: 300 tasks to train video models that reason natively in images
TheTuringPost · x · 2026-09-29
A large research group proposed VBVR-Pro, a suite of 300 tasks for "native visual reasoning" — using images and video as part of the reasoning process itself.
- Compared 30+ models with matched tasks: 1.25M training examples and 50 held-out task types; the goal is a "Verification Age" for video models — tasks with checkable outcomes used as RL feedback.
- Rule-based rewards lifted a score from 0.470 to 0.548, beating VLM-based rewards (0.508).
- Ablations show intermediate images matter: removing intermediate text dropped one model from 0.638 to 0.533; removing images crashed it to 0.099.
- After VBVR-Pro training, Wan2.2-I2V-A14B improved on external benchmarks: V-ReasonBench 10.21→38.22, VideoThinkBench 25.71→52.86, RULER-Bench 57.89→66.96. Video generation was strongest on continuous object tracking over time, and alternating text/images used less compute.
Related event: 52 Researchers Unveil VBVR-Pro Benchmark for Native Visual Reasoning(2 posts)→
More from Multimodal
- PrunaAI's P-Video-2 Pro models tie for #2 on Design Arena image-to-video leaderboard at Elo 1325 — guennemann · 2026-09-29
- QuiverAI's Arrow 2 Telos hits 1624 Elo, first model to break 1600 on SVG Arena leaderboard — stuffyokodraws · 2026-09-29
- Training FLUX.1 LoRAs on an 8GB RTX 5060: what optimizations work? — Wide_Director_8897 · 2026-09-29
- Flatbed debuts: an AI-native video editor where every asset is individually promptable — rchardkovacs · 2026-09-29
- One Image to a Walkable World: Hyper3D + GPT-6 Rebuilds Scenes in Three.js — Scobleizer · 2026-09-29
- Light Field Primitives: differentiable primitives replace dense ray databases for real-time novel view synthesis — zhenjun_zhao · 2026-09-29