VBVR-Pro: 300 visual reasoning tasks plus verifiable rewards for RL training
机器之心 · wechat · 2026-09-20
Researchers from NTU, CMU and Berkeley released VBVR-Pro, a scalable suite for native visual reasoning with papers, data, models and code all public:
- 300 visual reasoning tasks across perception, spatial, transformation, abstraction and knowledge, yielding 1.25M training samples rendered as both video and interleaved formats; trained models gain up to 20+ points on 7 unseen benchmarks.
- Verifiable scorers for 100 tasks reach >60% agreement with humans, beating GPT-5.5 (0.54) and Gemini-3.1-Pro (0.52), nearing the 0.77 human ceiling.
- Modality comparison over 30+ models: images suit end-state tasks, interleaved formats discrete states, video continuous dynamics; removing intermediate visual states hurts far more than removing text, revealing Chain-of-Step behavior.
- RL: the scorers work as verifiable rewards — Wan2.2-TI2V-5B improves from 0.470 to 0.503 with SFT and 0.548 with RLVR, with gains on out-of-domain tasks.
More from Multimodal
- User shows off Midjourney v8.2 portraits in 'My Different Sides' series — chrisfirst · 2026-09-20
- RefGarden turns one prompt into 100+ images and archive clips from NASA, The Met — round · 2026-09-20
- Qwen Image 2.1 open-source release counted down to hours away — CeFurkan · 2026-09-20
- Qwen Image 2.1 dropping tomorrow as PR merges into ComfyUI — fruesome · 2026-09-20
- CapoCut launches AI assistant for natural-language video editing — xiaohu · 2026-09-20
- LTX Video 2.5 Measured: 121-Second Single-Pass Generation, Resolution-Duration Equation Derived — Robotman2100 · 2026-09-20