FactoSR: RL with Factorized 4D Objectives Boosts VLM Spatial Reasoning
HKUST-GZ2 · hf · 2026-09-08
HKUST(GZ) releases FactoSR ("Unfold The World"), a method that improves vision-language models' spatial reasoning by decomposing 3D and temporal recovery into factorized geometric sub-objectives, each optimized via reinforcement learning. The project is available on Hugging Face.
More from Research
- New preprint: shared circuits predict LLM arithmetic generalization across formats — burny_tech · 2026-09-08
- Mathematician Andrew Snowden's new preprint used ChatGPT for many arguments — littmath · 2026-09-08
- 99.9% Accuracy Without Transparency Is Overfitting, Not an Edge — AryHHAry · 2026-09-08
- Meta's HumanCLAW benchmark: top VLMs fail badly at acting through a body across 1,218 tasks — wzenus · 2026-09-08
- Epoch: Astra beats most humans at card game but shows little continual learning — teortaxesTex · 2026-09-08
- Building World Models From Scratch, Part 1: Tokenizing Super Mario Land — Available_Pressure47 · 2026-09-08