3DZip: Efficient Token Compression Framework for 3D Question Answering

Changwoo Baek · hf · 2026-08-04

Current 3D vision-language models (3D VLMs) construct geometry-aware tokens by projecting 2D visual features into world coordinates, generating thousands of tokens per scene and incurring massive computational and memory overhead.

To address this, researchers propose 3DZip, a three-stage token compression framework:

Experiments on three 3D QA benchmarks show that 3DZip retains 94.7% of original performance with only 128 tokens, achieving a 1.92x speedup in inference.

Original post →

More from Research

Research channel →