FactoSR: RL with Factorized 4D Objectives Boosts VLM Spatial Reasoning

HKUST-GZ2 · hf · 2026-09-08

HKUST(GZ) releases FactoSR ("Unfold The World"), a method that improves vision-language models' spatial reasoning by decomposing 3D and temporal recovery into factorized geometric sub-objectives, each optimized via reinforcement learning. The project is available on Hugging Face.

Original post →

More from Research

Research channel →