Agentic SfM: Training MLLMs with RL Toolchains to Boost 3D Reconstruction

gabriberton · x · 2026-07-04

Researchers propose the concept of "Agentic SfM": training Multimodal Large Language Models (MLLMs) via reinforcement learning to use tools like image cropping, image matching, and COLMAP to solve 3D reconstruction challenges in complex scenes. Current image matching models perform poorly on low-overlap image pairs, whereas MLLMs (like QwenVL) can effectively distinguish between scenes with similar appearances (doppelgangers). Because the task is verifiable and supports curriculum learning sorted by difficulty, training costs are relatively low, making the overall approach highly feasible.

Related event: Agentic SfM: Training MLLM With RL Toolchain to Boost Hard-Case 3D Reconstruction(4 posts)→

Original post →

More from coding & agent

coding & agent channel →