VoxelTTO: voxel-aligned feed-forward 3DGS with test-time optimization, 80 GPU hours
zhenjun_zhao · x · 2026-09-22
VoxelTTO decodes Gaussians from a global voxel representation instead of pixel-aligned primitives, adapts frozen VFM via lightweight LoRA with pose supervision at test time, and uses stochastic solid volume rendering. Trained in 80 GPU hours, it improves RGB-D NVS and pose estimation on Replica, T&T and DTU.
More from Embodied
- Snap Specs' $2,195 Launch Faces 'Creep Glasses' Backlash and Apple's $1,999 Foldable iPhone — adariostrange · 2026-09-22
- Huawei MateBook Fold review: a 1.16kg 13-inch notebook that unfolds into an 18-inch 3.3K OLED, ~$3,700, China-only — anthara_ai · 2026-09-22
- Gemini beats GPT-6 Astra at robot capture the flag, winning 70% of matches — chris_j_paxton · 2026-09-22
- Meta Ray-Ban glasses talking to Muse could become the best consumer AI app overnight — ChrisUniverse · 2026-09-22
- JHU to host Scalable Tactile Sensing for Dexterous Manipulation workshop at IROS 2026 — _krishna_murthy · 2026-09-22
- SLIM-init: line-feature VIO initialization for degenerate motions, accepted to IROS 2026 — zhenjun_zhao · 2026-09-22