0.5B SpatialLM turns phone video into fully labeled 3D room reconstructions
joemeno · x · 2026-10-10
- A 0.5B model, SpatialLM, converts ordinary phone videos into fully labeled 3D room models, recognizing walls, doors, furniture, and overall layout.
- It was trained as a standard multimodal LLM, but unlike most vision models that see pixels, SpatialLM understands space.
- The open-source project by manycore-research was accepted at NeurIPS 2025 with 4.8k GitHub stars; SpatialLM1.1 (Llama-1B and Qwen-0.5B) doubles point cloud resolution with the Sonata encoder, and the dataset is on Hugging Face.
More from Models
- Gemini 4 Flash never announced, but community rumors are already flying — opmgyhx · 2026-10-10
- Theoretical physicist says ChatGPT Astra solved nearly all high-energy theory benchmark problems — burny_tech · 2026-10-10
- Several Juicy Frontier Models Expected Next Week, Evoking Strawberry-Era Hype — iruletheworldmo · 2026-10-10
- Unsloth shows fine-tuning Qwen into a decision model in ~2 min on one DGX Spark, 81% accuracy — danielhanchen · 2026-10-10
- What earns a new model a permanent spot in your agent? Devs weigh in after testing M3.1 Flash Preview — Willing_Amphibian825 · 2026-10-10
- vLLM Semantic Router: small-first routing cuts cost 54% while adding 16.45 accuracy points — vllm_project · 2026-10-10