OctLLM Uses Sparse Octree Tokens to Hit 3D SOTA While Preserving Language Ability
Ran Dan · hf · 2026-10-05
OctLLM addresses two limitations of existing 3D LLMs: latent codebook/coordinate-text representations that strip spatial structure, and backbone fine-tuning that erodes general language ability.
- Geometry enters as an explicit sequence of octree occupancy tokens; randomly emptying penultimate-level nodes yields a Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding.
- 3D capacity is added via separate trainable branches in a subset of blocks, keeping the frozen vision-language pathway intact, with the two streams interacting through shared self-attention.
- It sets a new SOTA among unified multimodal LLMs: image-to-3D FID down 17.4% and render-grounded captioning up 28.7 points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.
More from Multimodal
- User has Claude Fable 5.5 make a Rocket League edit — imjustnewatai · 2026-10-05
- One Claude chat turns a 50-second poem into an 18-scene AI film — aziz4ai · 2026-10-05
- EditHero: First Long-Horizon Part-Level 3D Editing Benchmark Shows LLM Agents Win — Ruihan Yu · 2026-10-05
- DEFINE Decouples Accent from Voice in Zero-Shot TTS, Matching Two-Model Cascades with One Model — amaai-lab · 2026-10-05
- Diptych: AI-Assisted Reference Listening Lets Musicians Define What to Compare in Music Production — amaai-lab · 2026-10-05
- FrameMorrow Selects History Frames by Predicted Future Needs, Boosting 11 Long-Video Generators — NationalUniversityofSingapore · 2026-10-05