VoxelTTO: voxel-aligned feed-forward 3DGS with test-time optimization, 80 GPU hours

zhenjun_zhao · x · 2026-09-22

VoxelTTO decodes Gaussians from a global voxel representation instead of pixel-aligned primitives, adapts frozen VFM via lightweight LoRA with pose supervision at test time, and uses stochastic solid volume rendering. Trained in 80 GPU hours, it improves RGB-D NVS and pose estimation on Replica, T&T and DTU.

Original post →

More from Embodied

Embodied channel →