SpaCeFormer: Interactive Open-Vocabulary 3D Instance Segmentation Without 2D Images

realChrisChoy · x · 2026-07-08

Releasing at ICML 2026, SpaCeFormer achieves open-vocabulary 3D instance segmentation by directly predicting labeled 3D instance masks from point clouds, eliminating the need for 2D image inputs or external bounding box generation models. With an inference speed of 0.1–0.3 seconds per scene, it supports interactive real-time use. While state-of-the-art methods rely on multi-view image streams and external models like YOLO/SAM, taking hundreds of seconds per scene, this method explores the performance boundaries of pure 3D end-to-end solutions.

Related event: SpaCeFormer: Open-Vocabulary 3D Instance Segmentation Without 2D Images Goes Open Source(2 posts)→

Original post →

More from Research

Research channel →