Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation

zhenjun_zhao · x · 2026-10-08

The MGE paper improves 3D foundation models by strategically dropping frame tokens from global attention during training and distilling from a pretrained full-context teacher. This yields richer per-frame geometric representations, much stronger performance under occlusion and doppelganger views without hurting standard benchmarks, and enables an Anchor-Guided Adaptive token merging technique for more efficient inference.

Original post →

More from Research

Research channel →