OctLLM encodes 3D geometry as explicit octree token sequences without sacrificing language ability
_akhaliq · x · 2026-10-05
OctLLM addresses two limitations of existing 3D LLMs: compressing shapes into latent codebooks or coordinate text removes spatial structure from what the model observes, and acquiring the 3D modality via fine-tuning overwrites the backbone's general language ability.
Key points:
- Geometry enters as an explicit 3D sequence of octree occupancy tokens
- Full octree sequences grow rapidly with depth, so OctLLM randomly empties penultimate-level nodes and omits descendants while preserving shape
- This yields a shorter, coordinate- and depth-anchored sparse sequence
Authors are from Peking University and collaborators (arXiv 2610.02388).
Related event: OctLLM unifies 3D and language with sparse octree tokens(2 posts)→
More from Multimodal
- User builds stunning interactive 3D Rolex movement explainer with Opus — tristanbob · 2026-10-05
- Creator shows off work made with Grok Imagine video generation — tetsuoai · 2026-10-05
- One Prompt Gets Fable to Generate a 5,000-Year History of China Video — FuSheng_0306 · 2026-10-05
- VNCCS 3.2.0 adds Qwen Image 2.1, native alpha sprites, and MiniMax H3 video model for sprite generation — AHEKOT · 2026-10-05
- UniMate Releases 3D Rigged Skeleton Models on Hugging Face, ComfyUI Integration Proposed — RazsterOxzine · 2026-10-05
- Full Making-of Released for AI Music Video 'Close All the Windows' — All Prompts and Orchestration Logs Included — Afinetheorem · 2026-10-05