Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation
zhenjun_zhao · x · 2026-10-08
The MGE paper improves 3D foundation models by strategically dropping frame tokens from global attention during training and distilling from a pretrained full-context teacher. This yields richer per-frame geometric representations, much stronger performance under occlusion and doppelganger views without hurting standard benchmarks, and enables an Anchor-Guided Adaptive token merging technique for more efficient inference.
More from Research
- NVIDIA's PivotOPD teaches agents to prevent and recover from pivotal early mistakes — NVIDIAAI · 2026-10-08
- ByteDance's DMAD hits FID 1.04 one-step on ImageNet, SOTA few-step visual generation — ByteDance · 2026-10-08
- BioinvestGPT reportedly predicts 5 of 6 major drug trial outcomes before results announced — Polymarket · 2026-10-08
- EigenDEXplore: better robot-hand exploration using coordinated human hand motion patterns — leto__jean · 2026-10-08
- Nathan Lambert publishes the essential open-models reading list — kevinsxu · 2026-10-08
- EMNLP 2026 paper RECAP trains reasoning models to recover from unsafe trajectories — pinyuchenTW · 2026-10-08