UniVR: A New Approach to Pure Visual Reasoning

ByteDance · hf · 2026-07-17

ByteDance proposed UniVR, exploring how to learn broader world knowledge relying solely on pure visual demonstrations, while simultaneously covering three capabilities:

The core method is VR-GRPO: a reinforcement learning paradigm combining global and step-level rewards to constrain logical and physical consistency during reasoning, without relying on task-specific heuristics or image-text paired data.

They also built the VR-X benchmark, aggregating 16 sources to cover long-horizon manipulation, spatial puzzles, and physical reasoning, serving as a comprehensive evaluation suite under a pure visual protocol. Experiments show UniVR improves performance on VR-X by up to 25%, and its enhanced visual reasoning capabilities also feed back into multiple multimodal understanding benchmarks. Code, data, and models are all open-sourced.

Original post →

More from Multimodal

Multimodal channel →