SpaCeFormer: Interactive Open-Vocabulary 3D Instance Segmentation Without 2D Images
realChrisChoy · x · 2026-07-08
Releasing at ICML 2026, SpaCeFormer achieves open-vocabulary 3D instance segmentation by directly predicting labeled 3D instance masks from point clouds, eliminating the need for 2D image inputs or external bounding box generation models. With an inference speed of 0.1–0.3 seconds per scene, it supports interactive real-time use. While state-of-the-art methods rely on multi-view image streams and external models like YOLO/SAM, taking hundreds of seconds per scene, this method explores the performance boundaries of pure 3D end-to-end solutions.
More from Research
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21
- LeRobot v0.6.0 adds end-to-end 3D depth training data for robots — RemiCadene · 2026-07-21
- Multiagent v2 playbook calls for 64 agents, diverse proof routes and adversarial checks — danshipper · 2026-07-21
- HarmonicMath says Lean autonomously solved eight previously studied open problems — MarioKrenn6240 · 2026-07-21
- SeeSE3 finds 3D structure emerging in frozen vision features and camera-pose alignment — ducha_aiki · 2026-07-21