NTU, CMU, Berkeley release VBVR-Pro: 300 verifiable tasks for native visual reasoning
jiqizhixin · x · 2026-09-29
Native visual reasoning is gaining attention: language is no longer the only path, as models can directly perceive, explore, and update visual space. Three questions block progress: what tasks should train visual reasoning, how to judge whether generated images/videos are correct, and which modality — image, video, or interleaved text-image — fits best. Without verifiable rewards, training loops lack reliable signal.
NTU, CMU, Berkeley and collaborators present VBVR-Pro (A Scalable and Verifiable Suite for Native Visual Reasoning), fully open-sourcing paper, data, models, and code. The suite defines 300 visual reasoning tasks across abilities and complexity, builds a benchmark from 100 of them with a verifiable scorer per task, and constructs 1.25 million high-quality training samples.
Related event: Researchers unveil VBVR-Pro benchmark for native visual reasoning(3 posts)→
More from Multimodal
- Physics-simulated instruments played by an LLM: full rehearsal finally succeeds — alexanderchen · 2026-09-29
- AI video tip: import a movie scene's camera move, replace everything else — LearnWithBishal · 2026-09-29
- AMD to acquire Fei-Fei Li's World Labs for $8.2 billion — pstAsiatech · 2026-09-29
- Villain-boss POV video made with Magnific and Seedance 2.5, prompt shared — charis_ai · 2026-09-29
- Ask HN-adjacent: local faceswap workflows all broken after reinstall — DwemerNose · 2026-09-29
- PrunaAI's P-Video-2 Pro models tie for #2 on Design Arena image-to-video leaderboard at Elo 1325 — guennemann · 2026-09-29