One model unifies four 3D part segmentation tasks, beating specialized baselines by up to 34.5 mIoU

PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

Zhe Zhu, Yiheng Zhang, Peng Li, Zixing Zhao, Honghua Chen, Yaqing Zhang, Le Wan, Zhiyang Dou, Cheng Lin, Yuan Liu, Mingqiang Wei, Wenping Wang

SIGGRAPH Asia 2026 (ACM Transact

cs.CV, cs.GR

2026-09-22

PartLLM casts 3D part segmentation as autoregressive generation: one MLLM handles text-guided, interactive, and full-shape settings, beating prior baselines by up to 34.5 mIoU.

What problem this solves

3D part segmentation splits a mesh or point cloud into semantic parts (chair back, sword hilt) and sits upstream of asset editing, rigging, and robotic grasping. The field is splintered. PartField-style methods output unlabeled geometric pieces; FIND3D and CoSMo3D ground parts from user-supplied text but never decide how the whole object decomposes; PartSAM and P3-SAM handle point clicks with equally nameless output. Table 1 lines up eight representative methods and none covers point interaction, text grounding, category-agnostic decomposition, and semantic full-shape segmentation at once.

The split wastes more than convenience. Different datasets observe the same part organization from different angles: some label only queried parts, some provide unlabeled regions, some annotate complete decompositions, and training a separate model on each keeps these signals isolated. The paper goes one step further: part segmentation is inherently ambiguous, since one sword can be split two ways or into a dozen parts depending on intent, which makes it a conditional generative problem rather than a discriminative one with a fixed answer.

Method

PartLLM is built on Qwen3VL-4B and runs in two stages: the MLLM first says which parts to carve out, then a decoder assigns every point to one of them.

Training combines a next-token loss with point-level cross-entropy over the dynamic label space, end-to-end for 5 epochs on 48 H20 GPUs. 323K shapes (PartNeXt, 3DCoMPaT200, PartVerse, 75K MLLM-annotated licensed assets, plus the unlabeled HY3D-Bench subset) expand into 1.1M samples in one shared format, which is the data dividend of unification.

Results

SettingBenchmarkBest baselinePartLLM
Full-shape (category-agnostic)PartNeXtPartField 42.756.7
Full-shape (category-agnostic)3DCoMPaT200 coarseSAMPart3D 51.686.1
Full-shape (category-agnostic)HY3D-BenchPartSAM 40.568.3
Text-guidedPartNeXtCoSMo3D 37.978.7
Text-guided3DCoMPaT200 fineCoSMo3D 26.180.0
Interactivemulti-round—up to +10.1 at one click, +18.5 at seven

The 3DCoMPaT200 margins are the headline: +34.5 mIoU over the strongest baseline at coarse granularity, +32.0 at fine. Number control works too: asking for 2, 5, or 10 parts yields progressively finer decompositions that stay semantically coherent. Semantic full-shape segmentation (masks must carry correct names) had no ready-made competitor, so the authors bolted an MLLM naming stage onto PartSAM and P3-SAM; those two-stage baselines lose 17.5 to 23.6 points under semantic matching, while PartLLM loses 11.1 to 14.2, a sign its regions are already tied to the names it generates. Real scans hold up: 72.2 vs PartSAM's 55.3 on FAUST full-shape, and 54.8 vs PatchAlign3D's 35.6 for text-guided segmentation on AKBSeg.

The most informative ablation swaps the joint softmax back to SAM-style independent binary masks: text and interactive settings dip only slightly, but full-shape mIoU falls from 74.2 to 62.9, so competitive assignment matters most where parts are most numerous. Removing the coarse 3D box from each hypothesis hurts most of all, dropping text-guided from 80.0 to 55.4; names alone cannot localize parts. The data-composition ablation shows the four supervision types transfer to one another, with full-shape mIoU climbing from 56.2 to 74.2 as sources are added. Across ten sampling seeds the model scores 78.5 ± 2.2 on PartObjaverse-Tiny, and test-time random rotations cost almost nothing (78.8 vs 78.3).

Why it matters

For 3D content pipelines, one model replaces four specialized ones, with open-vocabulary names, granularity set by a sentence, and masks that plug straight into editing (the paper demos material swaps and structural edits). The bigger signal is methodological: once heterogeneous annotations collapse into one prompt-response format, task diversity itself becomes a scaling axis, the curve is still rising at 1.2M samples, and even unlabeled geometry data benefits every setting. This is consolidation at the task-definition level, not a point or two shaved off a benchmark. The bill is real too: a 48-GPU training budget, and inference latency that grows with part count (competitive at moderate counts, per the paper's own timing).

Limitations

The authors' own failure case: when the object category is misread from geometry alone, the whole decomposition follows it. A router mistaken for a bed gets hallucinated bed substructure; adding the category to the prompt fixes it. For a generation-first architecture this failure mode is structural.

A few things to weigh. The semantic-aware metric (SA-mIoU) depends on GPT-5.5 judging whether predicted and ground-truth part names are equivalent; the matcher is independent of the Qwen backbone, but the matching boundary is still set by another LLM, and open-vocabulary naming is exactly where PartLLM gains. Semantic full-shape segmentation had no comparable baseline, so the two-stage competitors were constructed by the authors. The 75K auto-annotated training assets get no separate quality evaluation. In the interactive comparison only some baselines run the multi-round protocol while others report one-click numbers. Real-world evaluation covers two small scan benchmarks, FAUST and AKBSeg.

Terms

Source

What people are saying

Related papers

All paper explainers