EPFL paper recasts multi-view stereo as seq2seq, beating MVS and feed-forward baselines
CSProfKGD · x · 2026-10-03
A NeurIPS paper from Pascal Fua's group at EPFL shows feed-forward models distort geometry even with ground-truth camera poses. The authors reformulate multi-view stereo as a sequence-to-sequence task: a camera-aware transformer injects camera parameters via ray-map embeddings and uses a unified global cost volume to jointly predict geometry for all views. It achieves SOTA on public benchmarks, surpassing both MVS and FF reconstruction baselines.
More from Research
- Open Pretraining Run Matches Llama 3.2 1B at ~90% Lower Cost per Token — jon_durbin · 2026-10-04
- Neuralink pretrained AI on 50,000 hours of brain activity, hits 11.32 bps cursor-control record — mark_k · 2026-10-04
- Training a 128k-param neural cellular automata to 'see around corners' with sound — yacineMTB · 2026-10-04
- Watermarking proteins is easy — removing them with proteinmpnn is easier, researcher warns — anshulkundaje · 2026-10-04
- DNA sequence watermarks are 'more theater than security', says Stanford's Anshul Kundaje — anshulkundaje · 2026-10-04
- Meta paper: RL post-training hurts test-time scalability — the 'Sharpening Tax' — dair_ai · 2026-10-04