OctLLM encodes 3D geometry as explicit octree token sequences without sacrificing language ability

_akhaliq · x · 2026-10-05

OctLLM addresses two limitations of existing 3D LLMs: compressing shapes into latent codebooks or coordinate text removes spatial structure from what the model observes, and acquiring the 3D modality via fine-tuning overwrites the backbone's general language ability.

Key points:

Authors are from Peking University and collaborators (arXiv 2610.02388).

Related event: OctLLM unifies 3D and language with sparse octree tokens(2 posts)→

Original post →

More from Multimodal

Multimodal channel →